Fetching the latest programs, projects, and workspace data.
Find open source projects actively accepting contributors. Search repositories, filter by program milestones, difficulty tags, or tech stack.
Use our Orbit AI Matcher to find out! Get instant matching scores based on your developer skills, preferred frameworks, and contribution experience.
Convert your selected open-source project into a winning GSoC, LFX, or Outreachy application using Proposal Studio.
<p>General-purpose statistical modeling tools such as OpenBugs and Stan allow using easy-to-build modular likelihood models to fit models to data through MCMC or other techniques. Crucial to the efficiency of such tools, especially as parameter dimensions increase, are tools like HMC, which allow use differentiation and/or conjugacy for continuous parameters, so that fitting is not solely a matter of getting lucky with random Metropolis-Hastings proposals. Yet some models, especially nonparametric ones, also require discrete parameters and/or hand-written jump steps; and existing tools for HMC generally don't play well with discrete parameters. My project is intended to generally improve the Julia toolchain for MCMC when there is a mix of continuous and discrete parameters. My work will be based on Mamba.jl; though when I can make something that will be useful to Klara.jl too, even better. To make that plan concrete, I'll be implementing CrossCat, a general-purpose nonparametric model for "medium-size" tabular data including missingness and various data types. (Roughly speaking, "medium-size" might mean p<1e3, n<1e5).</p>
<p>Brian is a spiking neural networks simulator, which provides desirable syntactic sugar and flexibility to allow a wide variety of models without compromising rapidity. With a plethora of different sets of governing equations, neuronal models with complex biophysical properties, synapses with plasticity and parameters, a mechanism to derive generalized platform-independent model description becomes inevitable. Further, the model description mechanism helps in easy access and reproduction of the models, and thereby exasperating reference to the platform-specific source code shall be avoided. Currently, Brian uses <code>brian2tools.nmlexport</code> package to elegantly export Brian models to <code>NeuroML</code>. However, it is confined and cannot be extended easily to various other generalized model descriptors. Therefore, the project proposes an idea to create a generalized basic framework, which can coherently describe Brian models in a standard format and shall be easily extended to various other formats. The proposed idea would substantially enhance the interfacing functionality of Brian models with other standard model descriptors and thereby helps various users and research communities.</p>
<p>The goal of NASA’s NextGen research is to accommodate the traffic increase coming over the 15 years. One requirement for NextGen is to provide a better, at least the same, level of safety system, compared to the current ATS (Air transportation System). Using the Brahms Multi Agent System Modeling framework, researchers at NASA and the FAA are building models describing actions of flight and ground staff. These models are being explored to find better ways to coordinate to keep planes in the air, as well as how to land them quicker and safer.</p> <p>Within those models, the Java Pathfinder tool has been used to verify properties related to safety and workload. This verification process suffers from the (in)famous state space explosion problem: for the large model it is not possible to do complete state space exploration.</p> <p>In this proposal we propose to use the GALE optimizer (GALE is a fast heuristic algorithm which can find the optimal result to the multi-objectives problem by pruning the unpromising subspace) to find the most promising sub-state space in the models so that the verification process can be targeted and more efficient.</p>
<p>Brains on Board (<a href="http://brainsonboard.co.uk" target="_blank">http://brainsonboard.co.uk</a>) is an inter-university project with the aim of developing an autonomous flying robot with the learning abilities of a honeybee, with all computation performed on board. Part of the project involves developing a lightweight C++ library for running neural simulations in real time on small robots (<a href="https://github.com/BrainsOnBoard/bob_robotics" target="_blank">https://github.com/BrainsOnBoard/bob_robotics</a>). The BoB robotics framework currently lacks the following:</p> <ul> <li>A physics-based simulation engine</li> <li>Models of the sensors being used</li> </ul> <p>This project aims to integrate BoB robotics with a more sophisticated 3D simulation package i.e. Gazebo. The main contributions of this project shall be:</p> <ul> <li>Prepare Gazebo world environments for real-world 3D models that are currently available with bob_robotics. </li> <li>Modelling the sensors (specifically those that mimic the vision of insects) and robots used by the BoB robotics team into Gazebo.</li> <li>Integrating Gazebo with BoB robotics so that the simulated robots can be controlled using BoB robotics models, e.g. test out a navigation model on a good simulated model of a drone.</li> </ul>
With the advent of machine learning, statistical learning and modelling becoming more prevalent in today’s world, we can estimate values for an unknown data given a trained model i.e. finding value of data(D) given parameter (Θ) .Now what about the reverse? What if we want to find optimal parameter value, given a particular set of dataset. Why is this even required? The very simple answer to this can be found in every modelling and computation textbook. We cannot necessarily measure the other parameters. Let us take an example: There are several parameters while defining a neuron whose measurement is difficult to make, for example the axial resistance, the capacitance of the soma etc. Thus to achieve this process we make use of 2 tools, the first is the neuron library which is used for designing your neuron model and setting your parameters ,secondly is a statistical measurement tool used for the parameter estimation .This is done by checking for the voltage value generated by the neuron model and what might the parameter values be for that particular value. The reason to utilize Bayesian inference for parameter estimation is due to the specification of the prior data ,which will help us to increase the accuracy of our model which predicts approximate parameter values. Now the user could specify the prior data on the basis of any previous study or research material. Neuroptimus is an open source tool created for the same purpose. It is an interactive tool where the user can drop a dataset file. Train it and use it for parameter estimation . I have been tasked to utilize the prior ,the noise (for the calculation of likelihood),and the evidence to find the posterior distribution of the parameter I want to estimate. This is the main gist of the project, but there are various hurdles such as what if it is difficult to find the likelihood ,type of noise, computational costs for multiple parameters etc. These are some other things I believe I will be working on.
The project aims to improve longitudinal normative modelling in PCNtoolkit by validating, debugging, and extending velocity-based models to support multiple timepoints per subject. The work will include integrating the updated models into the toolkit’s modular architecture, adding tests and documentation, and developing functionality for predicting individual brain changes over time from multiple past observations.
This project will improve the RenAIssance OCR pipeline for 17th-century Spanish historical documents by strengthening preprocessing, removing marginalia, optimizing TrOCR inference, and adding a flexible correction and deployment stack. The plan is to convert raw scanned pages into cleaner, layout-aware inputs, improve printed-text OCR, extend the system to handwritten OCR with the available handwritten datasets, and integrate local LLM/VLM-based correction with external fallback support when needed. The final deliverables will be a modular OCR pipeline, a deployed Flask application, benchmarking against existing OCR tools, and a paper-style technical report with reproducible documentation.
<p>Deep neural networks are not yet being used (at scale) on microscopic images for research. With the right model architecture and training approaches, it is possible to train robust models which would help in the research efforts of many. These models can be used for:</p> <ul> <li>Segmentation and analysis of the C. elegans embryo</li> <li>Generating more image data with Generative Adversarial Networks</li> <li>Gain inferences from numerical metadata</li> <li>Large scale classification of images</li> </ul>
<p>Bridge the gap between Python community and the Julia community for the state of the art natural language processing models.</p>
<p>Gaussian processes (GPs) are flexible probabilistic models with a wide a variety of applications. However, in their simplest form they scale notoriously poorly with large datasets. There has been much work in recent years to develop approximate methods which enable GPs to be used with much larger datasets.</p> <p>This project will focus on implementing such approximate methods for Gaussian processes within the JuliaGaussianProcesses ecosystem with the goal of achieving performance which is competitive with similar established Python packages (specifically GPFlow and GPyTorch).</p>
<p>The project Conversion of large scale cortical models into PyNN/NeuroML involves the conversion of published large scale network models into open, simulator independent and testing them across multiple simulator implementations. In the previous edition of GSOC the large scale network model for the macaque cortex, proposed by Mejias et. al, was successfully converted. A natural extension of this model was proposed in a paper by Joglekar et. al . Instead of using non-linear firing rate models, the cortical area was simulated as a spiking neuronal network. This was extremely useful to investigate the propagation of activity in the synchronous and asynchronous regime of the network. My goal in this project is to convert this model to PyNN allowing the simulation in several simulators. As a secondary goal in this project, I would like to convert the model proposed by Demirtas et. al. that is a large-scale circuit model of human cortex incorporating regional heterogeneity in microcircuit properties inferred from magnetic resonance imaging (MRI) for parametrization across the cortical hierarchy and fitting models to resting-state functional connectivity.</p>
This project aims to significantly enhance the DeepChem library by integrating PyTorch Lightning to enable efficient multi-GPU training. Currently, DeepChem is primarily limited to single-GPU use, hindering work with large models and datasets common in scientific machine learning. The project will introduce wrapper classes for DeepChem's datasets and PyTorch-based models, making them compatible with Lightning's distributed training paradigms and also with the base classes of deepchem. This will provide users with a seamless way to scale their training across multiple GPUs, drastically reducing computation time and enabling research on more complex problems within the DeepChem ecosystem.
PyGeNN is a Python package for simulating spiking neural networks using GPUs. For this project, I will construct and simulate a PyGeNN model of the mouse primary visual cortex (V1) by reproducing the point-neuron V1 model created by Billeh et al. (2020). Taking advantage of GPU-acceleration, the PyGeNN model will be able to simulate V1 neural activity in real time. I will first implement a generalized leaky integrate-and-fire (GLIF) neuron class, which will govern the spiking behavior of neurons during simulation. Next, I will construct the network by connecting V1, lateral geniculate nucleus (LGN), and background (BKG) neurons, using the parameters calculated by Billeh et al. (2020). To validate the model, I will compare the results of PyGeNN simulations to those of BMTK/NEST simulations. Finally, I will construct an application that simulates V1 activity in response to a webcam stream in real time. If time permits, I will enhance this application by implementing a GPU-accelerated implementation of the filternet LGN model, which will increase the maximum frame rate that can be simulated in real time. The results of this project will include: (1) a complete model of mouse V1, (2) validation results and benchmark comparisons to BMTK/NEST simulations, (3) a webcam application showing simulated neural activity in real time, and (4) if time permits, a GPU-accelerated implementation of the filternet model.
<p><a href="https://github.com/JuliaText" target="_blank">JuliaText</a> is the JuliaLang organization that provides with packages to work with text. It currently lacks support for basic problems like Named Entity Recognition, Part-of-Speech Tagging, Dependency Parsing etc. which help serves as the basis for various language processing problems and analysing text.</p> <p>I propose to implement practical models for Named Entity Recognition and Part-of-Speech Tagging in Julia and extensively test and validate them. Robust and well-tested APIs for these two tasks will be written.</p>
<p>I propose to convert to NeuroML, test and validate a recent, full scale, multicompartmental CA1 hippocampus model by M. Bezaire <a href="https://elifesciences.org/content/5/e18566" target="_blank">DOI: 10.7554/eLife.18566</a>. By converting the model and putting it to Open Source Brain we would not only give an open source baseline models to scientist who aim to create and compare similar models, but also give a detailed step-by-step conversion guideline for the ones interested in using NeuroML.</p>
A Collective Cognition Model to serve as automated guidelines to open source communities. The model will project the activities, interrelationships, to keep the community self-sustaining and to help people stay in the community and become contributors. This will be validated by the data generate by an Agent based model and will serve as a verification component.
Computational models, based on detailed neuroanatomical and electrophysiological data, are heavily used as an aid for understanding the brain. An increasing number of studies have been published over the past, but remain available in simulator-specific formats. NeuroML is a standardized data format to describe the biophysics, anatomy, and network architecture of neuronal systems at multiple scales in an XML-based language. Hence, to improve accessibility, exchange, and transparency in the research of neuronal models, this project aimed to make neuronal mouse models in the Allen cell types database available in a simulator-independent format, using NeuroML.
<p>Different computational models have been developed in neuroscience to simulate neural systems, however, these models often use different programming languages, tools and techniques making it difficult to share and reproduce them among different research groups.</p> <p>NeuroML and LEMS have been introduced to standardise the structural description and dynamics of concepts such as ion channels and synapses in a machine-readable format, making computational models more reproducible, accessible and shareable among researchers. PyNN is a Python package that offers a common interface for different neuronal simulators in the field. Tying all these together, a curated database of neuronal models is publicly available to the community at the Open Source Brain (OSB) repository.</p> <p>This project focuses on the implementation of two Neural Mass Models (NMMs) using NeuroML/LEMS and PyNN. To validate and test the implementation of population models on NeuroML/LEMS, we will implement two previously published models and share them on the OSB platform.</p>
Biologically detailed models are essential tools in neuroscience, and automated methods are frequently used to build and validate such models based on experimental data. The open-source parameter optimization software Neuroptimus was developed to enable easy application of advanced parameter optimization methods, such as evolutionary algorithms and swarm intelligence, to various problems in neuronal modeling. Neuroptimus includes a graphical user interface, and works on various platforms, including PCs and supercomputers. While Neuroptimus uses various built-in cost functions and the eFEL feature extraction library to compare model behavior to experimental data, it severely limits the range of neuronal behaviors that can be targeted by optimization. However, the popular model-testing framework SciUnit allows for the implementation of tests that quantitatively evaluate arbitrary model behaviors. Therefore, the goal of this project is to extend Neuroptimus to use HippoUnit, an open-source neuronal test suite based on SciUnit, to evaluate model performance during optimization. This project aims to develop a seamless integration between HippoUnit and Neuroptimus, enabling the construction of detailed biophysical models of hippocampal neurons. The new Neuroptimus-HippoUnit integration will allow the optimization of a broader range of neuronal behaviors, which can ultimately lead to improved biophysical models of hippocampal neurons. Deliverables: 1- Implementation of an interface between HippoUnit and the new version of Neuroptimus to enable the optimization of a broader range of neuronal behaviors. 2- Extend the current supported tests to include all HippoUnit tests. 3- Integrate the HippoUnit with the Neuroptimus’s GUI so that users can run optimizations directly from the interface.
The nematode Caenorhabditis elegans serves as a pivotal model organism in developmental biology due to its well-defined cell lineage, compact genome, and the availability of comprehensive datasets detailing its development from a single cell to a complex multicellular organism. This project aims to leverage the existing DevoGraph framework to model the intricate developmental processes of C. elegans, specifically focusing on embryogenesis and the formation of its connectome, using advanced temporal graph techniques. The core methodology will involve employing Discrete-Time Dynamic Graph (DTDG) models, with an initial emphasis on the EvolveGCN architecture. The project will begin by constructing temporal graph representations from available C. elegans developmental datasets, including cell lineage and connectome data. Subsequently, EvolveGCN, a model known for its ability to handle evolving graph structures, will be applied and potentially adapted to capture the dynamic nature of these developmental processes. The expected outcomes include a functional implementation within the DevoGraph repository, a model capable of explaining key aspects of C. elegans development, and valuable insights into the efficacy of temporal graph techniques for addressing this complex biological problem. Furthermore, the project will explore potential future directions such as the visualization of the developmental graphs and the investigation of continuous-time models to capture finer developmental dynamics.
This project develops an open portal for goal-directed vision research, integrating eye-tracking datasets and machine learning models with tools for benchmarking and validation. A key focus is vision-language integration, enabling better task-driven attention modeling. This platform will also serve as a prototype for similar data+model initiatives on public hardware platforms, improving reproducibility and accessibility in multimodal AI research.
State-space models (SSMs) provide a flexible framework for modeling dynamic systems where latent states evolve over time. In econometrics, Dynamic Factor Models (DFMs) are widely used to capture the co-movement of multiple time series by assuming that a small number of latent factors drive the observed variables. The PyMC library already includes implemen- tations of SARIMAX, VARMAX, and structural state-space models, along with example notebooks for their usage. This project aims to extend the existing PyMC state-space module by implementing Dynamic Factor Models, aligning with the functionality available in statsmodels. Our implementation will follow a Bayesian approach, leveraging PyMC’s probabilistic programming framework and PyTensor for computation. Additionally, an accompanying example notebook will demonstrate estimation, forecasting, and causal analysis with the new model, ensuring accessibility for users. This enhancement will significantly expand PyMC's capabilities, empowering student, practitioners and researchers in econometrics, macroeconomics, and beyond to model complex time-series data probabilistically.
Biologically detailed models are useful tools in neuroscience, and automated methods are now routinely applied to construct and validate such models based on the relevant experimental data. The open-source parameter fitting software Neuroptimus (formerly Optimizer) was developed to enable the straightforward application of advanced parameter optimization methods (such as evolutionary algorithms and swarm intelligence) to various problems in neuronal modeling. Neuroptimus includes a graphical user interface, and works on various platforms including PCs and supercomputers. Neuroptimus currently uses various built-in cost functions and those implemented by the eFEL feature extraction library to compare the behavior of the models to the (experimental) target data. However, this approach severely limits the range of neuronal behaviors that can be targeted by the optimization. On the other hand, the popular model-testing framework SciUnit allows the implementation of tests that quantitatively evaluate arbitrary model behaviors. The aim of the current project is to extend the open-source neural parameter optimization tool Neuroptimus so that it is able to use test scores from the SciUnit framework as the cost function during optimization. Direct applications would include the construction of detailed biophysical models of hippocampal neurons using a combination of Neuroptimus and HippoUnit, an open-source neuronal test suite based on SciUnit. Deliverables: 1. Implement changes in the Neuroptimus core so that it is able to use SciUnit test scores as cost functions during optimization 2. Implement changes in Neuroptimus such that arbitrary combinations of built-in cost functions, eFEL features, and SciUnit test scores can be used 3. Implement changes in Neuroptimus such that the new features become accessible from the graphical user interface 4. Thoroughly test the newly implemented features and their integration with existing features of Neuroptimus
<p>Eclipse 4diac is an open source environment for programming distributed industrial automation solutions and control systems based on the IEC 61499 standard. One of the components of Eclipse 4diac is the 4diac IDE which is an integrated development environment for modeling distributed control applications compliant to the IEC 61499 standard. The current version of 4diac IDE does not support the validation of the models, which makes it difficult for users to get design-time feedback on inconsistencies regarding the models developed in 4diac. The Object Constraint Language (OCL) could be a solution to find issues in 4diac models since it provides capabilities for specifying generic constraints a model has to fulfill. The aim of this project is to develop OCL constraints and well-formedness rules to the metamodels of 4diac in order to improve the usability of the IDE.</p>