Fetching the latest programs, projects, and workspace data.
Find open source projects actively accepting contributors. Search repositories, filter by program milestones, difficulty tags, or tech stack.
Use our Orbit AI Matcher to find out! Get instant matching scores based on your developer skills, preferred frameworks, and contribution experience.
Convert your selected open-source project into a winning GSoC, LFX, or Outreachy application using Proposal Studio.
<p>MBDyn is a multibody dynamics solver which comes without any default graphical user interface for pre- and post-processing. This project is aimed at kick starting the development of an addon to a popular CAD software which can be used a pre-processor to MBDyn simulations. User can model the problem using the CAD software's python API which will be then transformed to a standard MBDyn input format file. This project is to start from scratch and is proposed to generate an MBDyn input code of at least a very basic CAD model, and code structure and documentations to help further developers contribute and complete the pre-processor.</p>
The interaction cross-section is an important quantity in high-energy physics, serving as a bridge between abstract theory and experiment. Is is also quite difficult to compute for physical processes involving many final-state particles. The SYMBA project utilizes machine learning, specifically a sequence-to-sequence (seq2seq) transformer model, to symbolically calculate a crucial component of the cross-section - the squared amplitude - from the amplitude for a given process. The goal of this project is to improve upon the current SYMBA results with a seq2seq transformer which takes advantage of the rich structure of Feynman Integrals, allowing one to directly input a Feynman Integral for a general process/theory and obtain an accurate symbolic representation of its squared amplitude.
The Mifos and Apache Fineract platforms, while offering comprehensive core banking solutions, present challenges in documentation due to their technical complexity and extensive domain-specific knowledge. To address this, the proposed project aims to leverage a class of large language models (LLMs) based on transformer architectures(especially reasoning based), to enhance documentation accessibility and user support. Building upon a preliminary LLM developed in 2024, the 2025 initiative will focus on scaling this system across the full documentation set, which involves scraping all the necessary sites,sublinks,slack channels and even youtube videos via their transcripts. This new project aims to completely revamp the way a normal Q&A chatbot does by leveraging state of the art models and latest techniques.
Open Climate Fix’s Quartz Solar forecasting model uses a simple rule-based "adjuster" that averages recent forecast errors to correct incoming solar forecasts. While effective in stable conditions, this method is rigid and might be suboptimal under rapid weather transitions or atypical error patterns. This project aims to compare this rule-based logic with TabPFN, a transformer model for tabular data. It seeks to evaluate whether TabPFN or its time-series variant, TabPFN-TS, can predict adjustments based on features such as time, forecast horizon, recent errors, and more. Through experiments, one key goal is also to determine whether TabPFN improves forecast skill—specifically the P50 quantile forecast—over the current approach, and whether it is efficient enough for deployment.
This project focuses on optimizing and deploying EduAid, an AI-powered tool that generates interactive quizzes from educational content. Key deliverables include: - Adaptive Difficulty Control: NLP-driven question generation targeting overlooked concepts using TF-IDF scoring and synonym replacement. - Multilingual Integration: Leveraging the T5 model to enable question translation into German, French, and Romanian. - Scalable Infrastructure: Deployment via Electron.js, Celery-Redis task queues for scaling by the factor of 10x. - Performance Enhancements: Model quantization (25% size reduction), ONNX runtime acceleration (65% latency improvement), and service workers for background processing. - Unified UI/UX: Revamped frontend with reusable components and IndexedDB migration for efficient data handling.
Graphite currently offers only a limited set of shape tools with basic gizmo support, limiting the scope of creative and precise editing. This project aims to enhance the Polygon Tool by introducing a variety of predefined shape nodes—such as Trapezoid, Star, Donut, Pie, Crescent, and more—each equipped with custom, interactive gizmos. To manage these features, a centralized Gizmo Manager will be developed to handle both general transform cages and shape-specific gizmo points. These gizmos will integrate with the Select Tool using a state-driven interaction model to enable intuitive, fine-grained shape manipulation. Each shape will expose unique transformation controls (e.g., adjusting edge curvature, radii, angles, and segment lengths), allowing users to craft complex designs efficiently.
DFlow is a tool designed to simplify the process of creating conversational models for virtual assistants. It aims to solve the problem of manual intent example creation in virtual assistant development. VAs require a set of examples of user intents to understand and respond to user queries correctly. However, manually creating these examples can be time consuming and vulnerable to errors, so the project has two objectives: 1. Develop a M2M transformation tool that can automatically generate intent examples from OpenAPI descriptions. 2. Automate the creation of intent examples from OpenAPI descriptions directly in the dflow platform. By achieving these objectives, dflow can make the VA development process much easier, saving time and effort for developers. The deliverables of this project include a functional M2M transformation tool and an automated intent example creation feature in the dflow platform.
This project focuses on extending the R interface of the torchvision library by implementing support for core computer vision datasets and model architectures. The scope includes adding native loaders for widely used datasets such as COCO and VOC, along with support for key models including Faster R-CNN, Mask R-CNN, FCN, Keypoint R-CNN, and quantized ResNet variants. The proposed work enhances the torch ecosystem in R by enabling advanced computer vision workflows with consistent APIs, GPU compatibility, and clean documentation. The implementation will follow package development best practices with robust testing, reusable components, and user-friendly examples. The resulting contributions aim to reduce the gap between R and Python in deep learning tools, enabling R users to build scalable, efficient, and modern vision pipelines.
The project addresses the rising energy consumption of ML workloads in high-energy physics (HEP) at CERN. While models like Baler and ATLAS top taggers are essential for real-time data processing at the LHC, they impose significant carbon and energy costs. This proposal introduces GreenML@CERN, a holistic framework for optimizing the energy footprint of scientific ML workflows. The solution combines cross-architecture energy profiling (GPUs, CPUs, TPUs), scheduler efficiency analysis (HTCondor, Kubernetes), algorithmic optimizations (quantization, pruning, dynamic training), and an automated Jupyter-based energy profiler with a real-time dashboard. Deliverables include energy-optimized versions of CERN ML models, benchmarking datasets, a profiling toolkit, and visual dashboards. The project extends CERN's ongoing GreenDIGIT sustainability efforts with developer-focused ML tools to help the community balance performance with energy awareness—contributing directly to CERN’s 2030 carbon neutrality goals.
For broadcasting messages, mailing lists are essential but high traffic lists frequently fill the inboxes with replies to threads of a particular topic that subscribers have no interest in. This project introduces Dynamic Sublists into the GNU Mailman 3 suite. This is an ideal feature for user-support community groups or for issue tracking where topics of interests/specific tasks diverge from the main discussion thread. The earlier rigid broadcast model is transformed into an opt-in thread model, ensuring only the root email of a new topic is broadcast to the main mailing list, after which the next round of replies will be routed strictly to a child list, delivering the mail to users who have explicitly opted in. Sublists are treated as dynamically made child MailingList objects linked via a parent_id. This project requires implementation across Mailman 3 Core, extending the SQLAlchemy IMailingList schema, implementing inheritance, building the email command workflow (-new to fork threads and standard -join/-leave aliases), and adding an admin specific toggle for Postorius.
<p>Synthesis is at the heart of organic chemistry. It is the process of constructing a molecule through a series of reactions. Efficient molecular synthesis is significant, especially in areas like medicinal chemistry. However, it can be challenging to plan a synthesis for a complex molecule.</p> <p>Retrosynthesis is an iterative process of breaking down a target molecule into a series of simpler molecules which are readily available. It provides a systematic way to plan a chemical synthesis. But, as the space of possible chemical reactions is vast, it is hard for chemists to make the right disconnections while optimizing for multiple factors such as the cost of precursors and efficiency of the reaction.</p> <p>At its heart, the challenge of making a single retrosynthesis prediction can be recast as a Machine Translation task. Consequently, Natural Language Processing techniques can be leveraged to tackle this using Deep Learning.</p> <p>This project is focused on extending the DeepChem Library to support Retrosynthesis models, specifically to make single step predictions. This will be achieved by implementing a Molecular Transformer model and enhancing support for Reaction datasets.</p>
The Ensembl Plants and Metazoa platforms face a significant metadata gap where critical biological context, such as ploidy, strain, and sex, is often documented in peer-reviewed literature but missing from formal INSDC sequence archives. This lack of structured metadata, particularly ploidy, creates technical bottlenecks in downstream comparative genomics and requires labor-intensive manual curation. To address this, I propose developing a standalone, Python-based "Retrieve-Extract-Verify" module that leverages Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) to automate metadata extraction from scientific papers. This tool will utilize literature APIs like Europe PMC, employing section-aware parsing and confidence scoring to ensure high data integrity with minimal human intervention. My previous experience in building Transformer-based bioprocess optimization models and genome-level embeddings directly informs my approach to handling complex biological datasets. The final outcome will be a ready-to-deploy Nextflow module designed for seamless integration into Ensembl’s production pipeline and provide the global research community with richer, more reliable genomic metadata.
Migrating content between different CMSs like Joomla, WordPress, and Drupal is challenging due to their distinct data structures. Traditional migration methods require custom tools for each CMS pair, leading to an O(n²) scalability problem. This project introduces a Common Content Model (CCM) as an intermediate format to simplify migrations to O(n) complexity. Using Model-Driven Engineering (MDE), we will dynamically adapt to various CMS schemas and automate migration script generation. The solution will be a Joomla component acting as a universal mediator to extract and transform content between CMSs. It can be installed on any Joomla instance, either directly or as a temporary migration app to handle API-based content migration between any CMSs. This project reduces development overhead, enables reusable migration structure, and positions Joomla as a central hub for content exchange across platforms. Deliverables: * CCM Joomla Component: A standalone migration component built for Joomla. * WordPress-to-Joomla Migration: Real implementation example using the CCM. * Testing Suite: Unit, integration, and E2E tests to ensure migration reliability. * Documentation: Clear guides for extending to other CMSs.
MP-BioPath uses rigorous mathematical programming to predict exactly how genetic mutations or drug perturbations cascade through Reactome's biological pathways. The current model needs better integration of tissue-specific realities and address the issue of artificially diluting predictive signals when multiple proteins share similar roles (entity-set dilution). AND While the Reactome team generated an expansion of these logic networks across their entire database, this massive architecture remains unvalidated at scale. This project aims to transform MP-BioPath into a scalable, context-aware predictive engine. I will architect an automated validation pipeline to stress-test the newly generated Reactome-wide logic networks, and then use this infrastructure to systematically benchmark MP-BioPath against traditional, topology-agnostic methods like Gene Set Enrichment Analysis (GSEA). I will build a transcriptomic-weighting module as the context later: by feeding real-world RNA-seq data into the optimization model, this extension will act as "live traffic updates" for the biological pathways, dynamically adjusting node weights based on actual tissue expression to resolve entity-set dilution and enable highly specific predictions.
At the end of Moore’s law, building domain-specific computer architectures is considered the next step in improving program efficiency. Hls4ml is an open-source project that translates machine learning (ML) models to high-level synthesis (HLS) code for deployment on hardware accelerators. The idea stems from the high-energy physics community at CERN. Google XLS (Accelerated Hardware Synthesis) is a novel framework that implements an HLS toolchain to produce synthesizable code for FPGA and ASIC applications. Thus, XLS can be integrated as one of the backends in hls4ml for transforming ML models into synthesizable code. The goal of this project is to integrate an XLS-based backend for the hls4ml framework. In doing so, hls4ml benefits from improved vendor compatibility and portability while also offering the potential to increase hardware efficiency. During development, we will benchmark performance and resource utilization against the existing backends (e.g., AMD Vivado, Intel Quartus). Furthermore, XLS provides a domain-specific language (DSL) called DSLX and an optimized compiler that can reduce compilation time. The expected results are an XLS backend prototype and benchmarks performed on various metrics for the synthesized code. Everything will be documented to facilitate use and future developments.
This proposal addresses the deployment gap in federated learning for genomics, where frameworks like FLAN enable distributed model training but lack standardization, security, and interoperability required for real-world biomedical environments. As a result, federated learning systems remain difficult to deploy across institutions and cannot scale in regulated settings. To solve this, the project transforms FLAN into a GA4GH-aligned, production-ready federated AI system by integrating key standards across the stack. It incorporates DRS for secure and standardized data access, TES/WES for portable and reproducible task and workflow execution, and TRS for containerized tool discovery. Security is strengthened using GA4GH Passports and Attested TLS for zero-trust, identity-aware communication, while Model Context Protocol (MCP) is introduced to enforce privacy constraints and execution policies across federated nodes. The deliverables include a refactored FLAN pipeline with DRS-based data access, modular training workflows executed via TES, end-to-end orchestration using WES (CWL/WDL/Nextflow), containerized environments registered with TRS, integrated Passport-based authentication and Attested TLS communication, a prototype MCP-based policy enforcement layer, and comprehensive documentation with a GA4GH-compliant reference implementation.
Proposal Summary: The project aims to achieve two primary objectives: efficient and precise particle identification and the development of effective lossless data compression techniques for the CMS Experiment at the Large Hadron Collider (LHC). Efficient and Precise Particle Identification: Traditional particle identification methods involve complex algorithms analyzing data from the CMS detector. Recent advancements suggest using convolutional neural networks (CNNs) directly on raw detector images, known as the "end-to-end" approach, for streamlined processing and competitive identification efficiency. The project seeks to explore and enhance vision transformers (ViTs) for particle reconstruction, benchmarking against established CNN models. Development of Effective Lossless Data Compression Techniques: The CMS experiment generates massive amounts of data from particle collisions, posing challenges in processing and storage. Data compression techniques are crucial for extracting relevant information while reducing data volume for efficient processing and analysis. The project aims to develop innovative masked auto-encoders (MAEs) with vision-transformer-based encoders for efficient data compression, enhancing downstream tasks' performance. Overall, the project endeavors to contribute to advancing particle physics research by improving particle identification accuracy and optimizing data handling through innovative data compression techniques.
Fragmented cultural heritage artifacts — such as Mayan stelae displaced from their original sites — exist today as scattered, incomplete pieces that are nearly impossible to physically reassemble. This project builds an AI-driven virtual reconstruction pipeline for Stela #43 from the Naranjo archaeological site, operating on 3D-scanned .PLY fragment data. The system combines a Point Transformer V2 backbone for geometric feature learning, a Siamese network for fracture surface matching, a Graph Transformer for globally consistent assembly reasoning, and a Trimmed ICP alignment module with pose graph optimization. A synthetic fracture dataset with realistic surface degradation simulation addresses the limited availability of real labeled data. Every learned component has a geometric fallback, and uncertainty quantification ensures low-confidence matches are flagged for human expert review rather than committed automatically. Deliverables include: a modular PyTorch reconstruction pipeline, a reusable synthetic fracture data generation tool, trained model checkpoints, a robust coarse-to-fine alignment module, a global pose optimization module, a full quantitative evaluation report comparing results against the existing geometric baseline, and complete technical documentation. The system is designed to augment archaeologist judgment — not replace it — by pre-screening thousands of potential fragment matches and providing interpretable, auditable match explanations.
<p>In GSoC 2020, The TFLite team and I released the <a href="https://docs.google.com/document/d/1zm9Ifiql4aFKTXMdEe0FXlEU83g6l5jWCO8a3BtHeI0/edit?usp=sharing" target="_blank">TensorFlow Lite Support Suite</a> consisting of TFLite Flutter Plugin, TFLite Flutter Helper Library, and a couple of example apps with tutorials. The project is receiving excellent comments from the community, as they are finally able to build performant Flutter ML apps with models and TF versions of their choice. The recent release of Flutter 2.0 should attract even more users. This year, we plan to extend the TFLite Flutter Support Suite based on the users' feedback and latest developments in the TensorFlow Lite ecosystem.</p> <p><a href="https://github.com/am15h/tflite_flutter_plugin" target="_blank">TFLite Flutter Plugin</a> will be improved for better handling errors and supporting more platforms along with Android and iOS.</p> <p><a href="https://github.com/am15h/tflite_flutter_helper" target="_blank">TFLite Flutter Helper</a> will be updated to match developments in its Java counterpart.</p> <p><strong>Flutter Task Library</strong> based on the TFLite Android Task Library to provide out-of-box support for popular use cases like detection, question, and answer, etc. The Flutter Library will leverage the performance and capabilities of the C++ Task APIs through dart ffi bindings. Example apps to demonstrate the usage of Task Library.</p>
This proposal bridges the gap between audio processing and machine learning by introducing an automatic differentiation (AD) library for the Faust programming language. This library empowers audio engineers to leverage machine learning within their familiar Faust environment. The core of the library will be pre-defined building blocks. These include Faust primitives with built-in derivative functions for common math and audio processing tasks. Additionally, there will be helper functions to simplify program construction and various loss functions specifically suited for audio processing. The proposal also explores Faust Neural Net Blocks (FNNBs), which are pre-written code modules representing specific neural network functionalities. These FNNBs will save development time, improve code maintainability, and allow audio engineers to focus on network architecture rather than low-level programming for each layer. The project deliverables will showcase the utility of the AD library in various phases. Phase I will deliver a working example of a differentiable filter or envelope in Faust. Phase II will focus on implementing a full-scale machine learning algorithm. This will be achieved through an autodiff file and a weights file. Finally, Phase III and IV will deliver an autoencoder implemented using FNNBs. By combining the AD library, FNNBs, and an architecture file specifying the training process, users can design and train complex neural network models for audio processing tasks entirely within Faust. This eliminates the need to switch between Faust and separate machine learning frameworks, while offering the benefits of automatic differentiation for efficient model training.
About 88 million U.S. households now identify as campers, and ~54 million households took a trip last year National parks run thousands of search-and-rescue operations every year (e.g., ~3,400 in 2022), and research shows tens of thousands of people were lost in parks over the 2004–2014 decade. First aid, navigation, shelter, and water skills reduce risk and help you protect others. That’s why I built Gemma Scout: an on-device, privacy-first AI assistant built for wilderness survival and outdoor adventures. Powered by a finetuned version of Gemma 3N, it combines domain-specific survival knowledge with real-time multimodal reasoning to guide users through critical tasks like plant identification, mushroom safety, shelter building, and first aid. The model was trained on over 30 million tokens and 7,618 curated Q&A pairs, sourced from expert handbooks and image datasets of edible plants and mushrooms. Leveraging Unsloth’s optimized finetuning framework with LoRA and quantization, Gemma Scout runs efficiently on mobile devices without internet access. Using LLM evaluations, Gemma Scout outperformed its base model in survival QA by a win rate of 89%, and achieved over 36% improvement in mushroom classification and 47% in plant identification. It supports multimodal chat via llama.cpp mtmd and Swift integration, enabling users to send images and receive grounded, expert-level responses—completely serverless. Whether you're camping, hiking, or lost in the wild, Gemma Scout helps you make smart, safe decisions—even when you're off the grid.
Apache Beam through its unified model for batch and streaming data-parallel processing pipelines, runners for executing them on a variety of distributed processing backends and ML specialized transforms within MLTransform (such as EmbeddingManager and other MLTransformProvider) make it uniquely positioned for building out RAG (Retrieval Augmented Generation) based applications. These applications are one of the most useful and commonly being built applications on LLMs (Large Language Models). For this project we will focus on building a knowledge base on a vector database for a text corpus, and enriching user's questions with matching text chunks using semantic search. This is a crucial part of any RAG applications and helps us in building the right prompt context for LLMs. We will implement the following deliverables to achieve this: 1. Build a Beam pipeline that takes in a batch text corpus from a public dataset as parameter to pipeline and uses MLTransform to generate and save Embeddings in batch mode to a vector database. - Initial scope: Wikipedia dataset with JinaAI Embeddings read from object storage and written to RedisIO to publish to known vector DB. 2. Build a Beam pipeline that takes in stream of text questions from clients and enriches it with related texts from the vector DB. - Initial scope: KafkaIO based reading of queries which is populated by producers independently and published in a different topic for results. 3. New enrichment handlers for vector database queries over Redis based Vector DB Stretch goals: 1. Implement enrichment handlers for OpenSearch (AWS supported)[4] 2. Implement enrichment handlers for Vertex AI Vector Search (GCP Supported)[5] The goal is to demonstrate semantic search building capabilities trivially using Beam and hence evaluation of search results is not tied to a broad benchmark (such as MTEB) for this project's scope.
Package Large Language Model (LLM) inference libraries, in particular vLLM. It is needless to explain how LLMs are important. Currently, the Debian archive only has PyTorch, but downstream applications are still missing. One of the most promising downstream applications is LLM inference. There are already people working on llama.cpp and Ollama, but vLLM still lacks lots of dependencies to land onto Debian. For multi-GPU inference and concurrency, vLLM has its advantages over llama.cpp. The missing packages are, for instance, transformers, huggingface-hub, etc. This project involves the Debian packaging work for vLLM and its dependencies that are missing from Debian, as well as fixing issues (if there are any) in existing packages to make vLLM work. Timelplan ~ 5/8: Start with the architecture-independent package and non-cuda variant. 5/8 ~ 6/1: I will start with the architecture-independent package and non-cuda variant. I already created the package for huggingface-hub so I think this is a good start and doesn't take much time. By this time, I will finish the package for tokenizers packaging. 6/1 ~ 6/30: Create the package for the transformers package. Then I will finish rest of the package for the vLLM package. 6/30 ~ 7/18: I will try to build the vLLM package with cuda variant. Now, I already checked that vllm works on CUDA 12.4 so I will try to build the package with CUDA 12.4 with a specific GPU(Maybe T4?). 7/18: Mid-term evaluation 7/19 ~ 7/30: Try to build the vLLM package with multiplatform. I will try to build the package with different CUDA versions and different architectures, including a multi-GPU environment. I will get the feedback from the mentor and continue to the next step. 8/1 ~ 8/25: I will try to finish the project as soon as possible and try to fix the issue. If I have some more time, I will continue with the SGLang packaging. 8/25 - 9/1: submit final work product.
<p>Jaeger is the industry-standard platform for distributed tracing. As microservice architectures grow complex, finding root causes in massive trace data becomes increasingly difficult. While Phase 1 of this initiative established a baseline AI assistant for natural language search, the system currently relies on hard-coded capabilities. This project (Phase 2\) aims to transform the Jaeger AI agent from a static chatbot into an extensible, user-programmable platform. The primary objective is to implement a "Self-Service Skills" framework, architecturally similar to "Claude Code Skills." This will allow end-users to teach the Jaeger AI new debugging workflows (e.g., "Analyze Critical Path" or "Detect N+1 Queries") by simply adding configuration files containing system prompts and logic rules, without needing to recompile the Jaeger binary. The applicant will build this extension within the Jaeger v2 (OpenTelemetry-based) architecture, utilizing **LangChainGo** to orchestrate interactions with Language Models (SLMs/LLMs). This project bridges the gap between generic AI reasoning and domain-specific observability expertise.</p><p><br></p><p>**Expected Outcome:**</p><p> - **Skills Engine Implementation:** An approach compatible with our [BYOA (bring your own agent)](https://docs.google.com/document/d/1qD0OpyRfq-JbO6MCB5gmVxsdPcpPhxz1R_pPnKDdYOg/edit?tab=t.0#heading=h.qgr5ifum0a9m) direction that dynamically discovers, validates, and loads user-defined "Skills" (prompts and tool definitions) from configuration.</p><p> - **Smart Analysis Features:** A polished implementation of Natural Language Search and Contextual Trace Explanation that intelligently leverages these loaded skills.</p><p> - **Local-First Support:** Verified compatibility with local model runners (e.g., Ollama, Llama.cpp) to ensure deterministic performance without sending data to public clouds.</p><p> - **UI Integration:** Enhancements to the Jaeger React UI to expose these AI capabilities and visualize the "reasoning steps" taken by the agent.</p><p> - **Documentation:** A complete guide for users on "How to Author Custom AI Skills for Jaeger."</p><p>- **Learning Opportunities:**</p><p> - **Agentic AI Architecture:** Learn to design stateful AI agents in Go that utilize "Tool Calling" and "Reasoning Loops" rather than simple text generation.</p><p> - **OpenTelemetry Internals:** Gain deep familiarity with the OpenTelemetry Collector architecture, as Jaeger v2 is built directly on top of it.</p><p> - **Cloud-Native Engineering:** Experience contributing to a graduated CNCF project, including navigating code reviews, writing design docs (RFDs), and adhering to open-source best practices.</p><p> - **Full-Stack Development:** Practical experience bridging a complex Go backend with a modern React frontend.</p><p><br></p>