Fetching the latest programs, projects, and workspace data.
Find open source projects actively accepting contributors. Search repositories, filter by program milestones, difficulty tags, or tech stack.
Use our Orbit AI Matcher to find out! Get instant matching scores based on your developer skills, preferred frameworks, and contribution experience.
Convert your selected open-source project into a winning GSoC, LFX, or Outreachy application using Proposal Studio.
This project aims to improve and populate the existing benchmark framework of PyBaMM. This will be achieved by - 1. Adding new benchmarks for prominent battery models (Single Particle Model, Doyle Fuller Newman Model, and Single Particle Model with electrolyte). 2. Adding user and developer documentation for benchmarking suite. 3. Creating new benchmarks for other models and PyBaMM's API. 4. Using files from ASV to visualize benchmarks locally. 5. Making the existing and the added benchmarks reproducible in any given environment. Deliverables - 1. Documentation for benchmarks in PyBaMM’s website 2. Template-like structure for benchmarking battery models available in PyBaMM. 3. Benchmarks for prominent models. 4. Benchmarks for other models and PyBaMM's API. 5. Scripts for automating and running benchmarks locally (stretch). 6. An environment in which all PyBaMM benchmarks can be reproduced (stretch).
This project aims to extend PySAL's spatial optimisation library (spopt) by implementing flow-based facility location models, specifically the Flow Refueling Location Model and the Deviated Flow Refueling Location Model. While spopt currently includes various node-based location optimisation models, it lacks flow-based models that are crucial for transportation infrastructure planning. The proposed implementation will include comprehensive data processing components for network and flow data, core optimisation model implementations with solver integration, and output processing tools for visualisation and analysis.
<p>The OpenCV's DNN Module allows us to run inference on a pre-trained Deep Neural Network in order to accomplish high end vision tasks with just a few lines of code. However OpenCV's model zoo needs some new additions. Models that have acquired State-of-the-art (SOTA) in different computer vision tasks are somewhat lacking in the model family that is currently listed in the OpenCV repository. This gives rise to the necessity of adding new, powerful models to the list. The goal of this project is to curate more models for ease of use by the OpenCV DNN module and put them in a place where they can be easily accessed such as LFS on Git. The proposed project will add six (potentially nine) new models to the OpenCV Model Zoo. All of the selected models have, either state-of-the-art results in the task they perform or belong to a task category not currently present in the model zoo (e.g. Image Generation). The proposed workflow and the established timeline take into account worst-case scenarios, guaranteeing the project's completion and minimizing risks.</p>
<h3>Adaptive Quantization based on an activity mask.</h3> <p>The human eye is more tolerant towards errors in areas of high activity and is quick to find out errors in areas of lower activity. To leverage this psychovisual characteristic, quantization can be made adaptive based on the activity mask. Activity masking will be implemented in two phases:</p> <ol> <li>Biasing the RDO based on the activity at a specific region.</li> <li>Varying the quantizer offsets across segments based on activity.</li> </ol> <h3>Optimizing the quantization algorithm using Trellis Quantization:</h3> <p>Using Trellis Quantization passing the activity measurements as the weights to the trellis. The output of this Activity masked Trellis is used for quantization. This helps the quantization perform better in PSNR metrics while increasing perceptual image quality significantly. This is a feature that has proved to perform better in the case of x264 and is expected to yield similar results at rav1e too.</p>
With the emergence of Internet of Things devices, lightweight models have become increasingly vital. This project focuses on building computationally efficient text segmentation models, enabling them to run effectively on devices with limited resources like memory, processing power, and battery life. I aim to improve the current model and expand support to 2-3 additional languages, potentially benefitting hundreds of millions more people. This will be achieved by restructuring the model using TensorFlow's Functional API, which facilitates the implementation of Conditional Random Fields. Additionally, I will be creating an intuitive model evaluation script to offer a customizable experience for selecting the most suitable model.
<p>This project aims to extend the functionality of pysal.spreg to deal with panel data econometric models. Spatial panels refer to data containing time series observations of a number of geographical units. Specifically, this project will focus on static panel models with fixed and random effects. Also, I will develop a framework to test the spatial lag, the spatial error model, and the spatial Durbin model against each other, as well as a framework to choose among fixed effects, random effects or a model without fixed/random effects. Thus, this project will handle panel data, estimate panel models and provide specification tests.</p>
<p>The aim of the project is to create an Active Deep Learning and Machine Learning model that would predict the class of the calls depending on the model the user selects. It would perform active learning while training by querying the user to annotate the data that the model has identified with lower confidence(other querying strategies would also be used). This data annotated by the user is again used for retraining the model. The prediction of any class by the model that the user desires would be based on the spectrograms generated by the audio files.</p>
<p>PyBaMM offers a way to compare new models by implementing models as expression trees that can be specified independently of the user's preference. This allows the model to be defined independently of the user's choice of parameters, spatial discretization, numerical methods and so on, which are plugged in during model processing. This project aims to use LaTeX to render the expression tree into a human-readable format using SymPy and also generate a file of the model equations to visualize the equations easily.</p>
Seeing the current increasing demand of high performance batteries, accurate thermal modelling of battery behaviour is essential. This project extends PyBaMM by adding the capability to simulate battery temperature in three dimensions. It starts by developing 3D meshes (box, cylinder etc) and FEM method. Then we develop a 3D thermal model for batteries, initially assuming a constant heat source. This thermal model is then coupled with existing battery models (like the SPM or DFN) so that the heat generated during battery operation can be simulated. The model will also eventually allow temperature data from the thermal simulation to feed back into the battery model, enabling more realistic, two-way interactions.
<p>The Women in Energy Open Source Mentorship Programme is a first-of-its-kind initiative designed to bring more women into open source energy modelling. Developed in partnership between Open Energy Transition (OET), LF Energy, the World Resources Institute (WRI), EnergyHaus Africa and the Centre for Net Zero, the programme builds structured, supported pathways that connect aspiring women energy modellers with experienced practitioners across the ecosystem. The programme addresses a critical gap, women remain significantly underrepresented in open source energy modelling communities worldwide. This is not just an equity issue. It is a capability gap. When the people building energy models do not reflect the diversity of the communities those models serve, the models themselves are weaker for it.</p><p><br></p>
Using Deep Regression techniques to decode Dark Matter with Strong Gravitational Lensing. Try to use the SOTA deep model (such as transformers) to do the regression task.
<p>Knowing about the progress and performance of a model, as we train them, could be very helpful in understanding it’s learning process and makes it easier to debug and optimize them. It could also help us affirm the results from our models or inspect them in case of counterintuitive behaviors. Hence, I aim to provide a step-by-step process of visualizing training statistics for people, who want to keep a tab on learning processes of their gensim models, and want to optimize it through experimentation with hyperparameters.</p> <p>The next phase of my project would aim to introduce and build a visualization module based on Gensim models and features. This would allow the interactive exploration of applications based on Gensim and would also enable users to do a qualitative assessment of their models and analyze the results. I aim to focus on implementing a visual framework for the exploration of topic models, taking cues from TMVE, pyLDAvis, and add more options to visualize data attributes and compare across the different topic models.</p> <p>This work can naturally be extended to various other features in Gensim and to the upcoming ones, to have an associated visualization.</p>
The objective is to develop a machine learning model to tag sound effects in streams (like police sirens in a news-stream) of Red Hen’s data. A single stream of data can contain multiple sound effects, so the model should be able to label them from a group of known sound effects like a Multi-label classification problem. The first step would be to develop a baseline model using existing pre-trained deep learning models and add to the Red Hen’s pipeline. Then the performance can be improved using transfer learning and fine tuning the existing model to achieve better accuracy. In this process, the models can be trained on sound effects from noisy or human labeled data sets after they are pre-processed to avoid acoustic domain mismatch problems.
DeepForest is a library for training and deploying deep learning models for forestry tasks. The library is built on top of the PyTorch deep learning framework and supports custom model architectures and loss functions, allowing users to create and train their models. It also supports multi-label classification, enabling the detection and classification of multiple tree species within the same image. The project I want to participate in is the Advancing Bird Detection and Classification in Hand-Held Airborne Imagery project and the goals are: 1 - A refined and optimized deepforest model for accurate bird detection and classification in hand-held plane imagery. 2 - An annotated dataset for training and testing, contributing to the improvement of the model's performance. 3 - Comparative analysis of bird count accuracy by species, providing valuable insights into the model's effectiveness.
<p>The GAS package for R aims to create an integrated computational environment to deal with Generalised Autoregressive Score (GAS) models. GAS models are particularly suited for univariate and multivariate time series and can be used to predict the evolution of several quantity of interest in economics and finance. Several research papers have demonstrated the superior ability of GAS models compared to alternative time series models. Most applications are in the fields of volatility modeling, systemic risk measurement, macroeconometrics, credit risk analysis and dependence modelling. The GAS package will offer the possibility to practitioners and academics to estimate GAS models using their data, perform predictions and simulations and obtain a graphical representation of the results. The data available at www.gasmodel.com will also be included for comparative and education purposes. The code will be written principally in C++ for computational purposes and then linked with R using the Rcpp, RcppArmadillo and RcppGSL packages. Code parallelisation will be included using the R parallel package and the OpenMP API. The package documentation will be written using the R roxygen2 package.</p>
<p>This project is to load 3D Models in PLY format from the CDLI database into the existing 3D Model Web Viewer using the Three.js library. The idea is to implement an RTI Viewer so that any CDLI tablet for which the PLY model is available has an easy-to-navigate option for the 3D viewer for that model.</p>
<p>Both astropy and Sherpa (<a href="https://github.com/sherpa/sherpa/" target="_blank">https://github.com/sherpa/sherpa/</a>) provide modeling and fitting capabilities; however, Sherpa’s features are way more advanced providing far more build-in models, a larger choice of optimizers and a real variety of fit statistics. Whereas Astropy's functional state-based interface is much more user friendly. Astropy is also more widely used. The main goal is the bring Sherpa’s optimizers and fit statistic functions to astropy. A more hopeful goal is to develop a bridge between both packages such that a user can use a astropy models completely interchangeably with Sherpa models and fitters. This will give astropy users access to fitting capabilities that required many years of developer time and that are unfeasible redevelop from scratch.</p> <p>Goals:</p> <ul> <li>Sherpa optimization routines embedded within astropy fitters and have them use astropy models.</li> <li>The ability to use of all sherpa's fit statistic functions while fitting.</li> <li>Have sherpa models imported through astropy act like astropy models.</li> </ul>
<p>The goal of this project is to build a new R package providing high-level user interfaces to many kinds of ecological models and implementing the analyses internally using NIMBLE. The project has two primary motivations. The first is to provide better tools for widely used ecological models. There are many individual R packages and other software that fit specific ecological models, but they lack flexibility in terms of the type of models and available Markov Chain Monte Carlo (MCMC) algorithms. The second, perhaps equally important, is to showcase how NIMBLE can serve as a computational engine “under the hood” of an R package, and to develop extensions to NIMBLE that will support its use in this way. The proposed package, called nimbleEcology, will provide a unified framework that allows users to model a variety of ecological process with a simple specification of the model structure, and estimate them using efficient and customizable MCMC samplers.</p>
In high-energy physics experiments such as those conducted at CERN's ATLAS project, the volume of data generated is immense, reaching up to 60 million MB per second. While lossless compression is already employed to manage this data, lossy compression—specifically of floating-point precision—offers more aggressive reductions, potentially decreasing file sizes by over 30%. However, this comes at the cost of irreversibly discarding information, raising the challenge of how to recover or approximate full-precision data for downstream analysis. This project proposes a novel approach using deep probabilistic models to reconstruct high-precision floating-point data from aggressively compressed representations, a problem coined as precision upsampling. The goal is to explore and compare the capabilities of three classes of generative models: autoencoders, diffusion models, and normalizing flows. Each model type offers distinct advantages: autoencoders are well-studied in neural compression, diffusion models are robust to noise and excel in reconstructing multi-scale structures, and normalizing flows offer exact likelihood estimation and invertible mappings that align well with the structured nature of physical data. Deliverables include a literature review on neural precision upsampling methods, a working baseline model adapted to CERN's PHYSLITE floating-point compressed data, and a comparative study of advanced probabilistic models—including variational autoencoders, diffusion models, and normalizing flows. Evaluation metrics will also be tailored to the data and implemented. Final outputs will include reproducible code, performance benchmarks, a written technical report, and optionally a publication or integration into the ATLAS codebase.
The {torchvision} library in mlverse is already near-feature-parity with the PyTorch library, but certain performance optimisation and state-of-the-art computer vision model support gaps need to be bridged. This project aims to bridge the performance optimisation and state-of-the-art computer vision model support gaps in the current mlverse library. The performance optimisation objective is to improve the performance of the library by moving computationally intensive functions, e.g., region proposal networks and non-maximum suppression, from R to C++. This will allow for seamless integration with libtorch, providing production-grade speed and performance for detection and segmentation tasks. The state-of-the-art computer vision model objective is to improve the library by adding YOLO, RT-DETR, and SAM models, among others, into the torchvision library. This will be achieved by providing end-to-end documentation of the models, e.g., using the R documentation system, roxygen. This project will result in improved performance by providing an R interface for performance-critical functions using Rcpp, as well as improved state-of-the-art computer vision models by adding YOLO, RT-DETR, and SAM models, among others, into the torchvision library. This will improve the performance, usability, and adoption of R as a deep learning library for computer vision tasks, making mlverse a competitive alternative
The project aims to build and publish python-cookiecutter templates for new PyBaMM-based projects. It is designed to simplify the setup process of the development environment for researchers and scientists interested in utilizing PyBaMM for battery modeling who may lack familiarity with managing Python environments or repositories. The project intends to enhance the accessibility and usability of PyBaMM for newbies and experienced users alike. Providing a standardized template with best practices and automation tools would lower the barrier to entry for adopting PyBaMM in battery modeling projects. This would in turn make battery modeling more accessible, efficient, and collaborative for the research community. The project also intends to implement "Model entry points'', allowing community contributors to create and share models of their repositories using the cookiecutter template without directly adding them upstream. This would not only let community contributors retain ownership and choose license terms but also grant flexibility to the PyBaMM team in supporting models, Including all of GitHub's functionality and infrastructure contained within the template.
<p>Super Resolution is a subset of algorithms that aim to up-sample a lower quality image to a higher quality one. It’s goal is to create an up-sampled copy that is as detailed and visually pleasing as possible. It is used in a wide range of fields, such as medical image processing or surveillance camera stream processing. Using deep learning models for super resolution is a widely researched area, as they generally achieve better accuracy than classical computer vision based algorithms. There are many popular type of models used, ranging from supervised learning to unsupervised learning methods. I propose to implement two deep learning based models to be a part of the OpenCV library. One is EDSR, which is a residual network based model, and is known for it’s high accuracy. The other one is LapSRN, which is a fast, but still accurate model, that can be deployed to real-time applications. So integrating these two OpenCV would obtain a model that achieves state-of-the-art accuracy, and another one that can be deployed on devices that have lower computational power, while still maintaining high accuracy.</p>
This project aims to modernize the Speak Activity by integrating a large language model (LLM) & small language model (SLM) to enhance the chatbot and integrate a modern TTS model to improve the voice features, making it more educational and engaging for early learners. The Speak Activity, traditionally used for promoting reading skills through synthetic speech, will be upgraded to provide a more natural and interactive experience. The primary objective is to implement an LLM-based chatbot offering improved conversational abilities and a more human-like interaction style. For the voice feature a high-quality, natural-sounding Text-to-Speech (TTS) model that accurately handles pronunciation and phonetics will be used. By fine-tuning the LLM to manage invented spelling and grammatical inaccuracies, the chatbot will offer a supportive and child-friendly interaction, guiding learners in an encouraging manner. The project will utilize Python and a fine-tuned model hosted on the cloud and accessed via an API endpoint. This design ensures compatibility and seamless integration with the existing Speak Activity, without adding extra dependencies. To support offline functionality, a lightweight, fine-tuned fallback model will be packaged with the activity. This fallback ensures continued usability even without an internet connection, albeit with slightly less accurate responses compared to the cloud-hosted model. This enhancement will make the Speak Activity a more effective and child-friendly tool, promoting better language learning through engaging and conversational interactions.
This project aims to develop a new beginner-friendly Machine Learning Library for Processing by (1) using the diverse model pool supported by Deep Java Library and (2) providing a user-friendly, step-by-step reference & tutorial page. A powerful Machine Learning library can expand the possibility of Processing by enabling the users to create cool AI-powered art, games, and interactive applications. Furthermore, a well-implemented ML library in Processing can be an effective tool to lower the barrier of AI/ML and teach ML to beginners in a friendly way. Though there are some ML libraries such as OpenCV for Processing and Deep Vision for Processing, there is no one universal ready-to-use ML library that is friendly for beginners. Additionally, the speech/audio models and text models are not supported by any of the existing libraries. By using Deep Java Library's extendable high-level API, this library resolves the aforementioned issues by enabling powerful pre-trained models from Deep Java Library's model zoo, including speech/audio and text models.