Fetching the latest programs, projects, and workspace data.
Find open source projects actively accepting contributors. Search repositories, filter by program milestones, difficulty tags, or tech stack.
Use our Orbit AI Matcher to find out! Get instant matching scores based on your developer skills, preferred frameworks, and contribution experience.
Convert your selected open-source project into a winning GSoC, LFX, or Outreachy application using Proposal Studio.
LLM is a hot topic, there are more and more frameworks to make the execution of LLM faster. WasmEdge already integrated the llama.cpp (https://github.com/ggerganov/llama.cpp) as one of the backend. Running LLM with CPU only is huge for those users who don't have GPU. We would like to integrate Intel Extension for Transformers (https://github.com/intel/intel-extension-for-transformers) as a new WASI-NN backend to provide a faster CPU inference performance. Expected Outcome: A new plugin provides a Intel Extension for Transformers WASI-NN (https://github.com/second-state/wasmedge-wasi-nn) backend, a test suite for validating the plugin, documents and examples for explaining how to use the plugin.
For now spopt has implemented several basic facility location models, providing the free open source for researcher, or organizations to use. However, there are still many useful improvements can be made. This proposal suggests improving the spopt package by implementing the P-Median model with Near-Far Cost Allocation. In the article by Church (2018), it proposed a new p-median model which can distinguish between near and far facilities, use both explicit and implicit variables for capacity allocations. Based on that, the writer plans to implement the p-median model with near-far cost allocation. By this, the spopt package can provide more accurate and efficient solutions to spatial optimization problems, typically for the problems with large data points, and large demand volume. This proposal will introduce the reason why the writer chooses this project, the technical details of this project, and the plan and delivery schedules.
<p>I will implement DeepFillv2 and add detailed Colab demos for it. The model will be "fine-tunable" to support the “reusable” characteristic of TensorFlow Hub.</p>
SQLancer randomly generates test cases consisting of potentially hundreds or thousands of SQL statements. In order to efficiently identify bugs in these statements, automatic reduction is needed to systematically remove the bug inducing statements until a minimal bug-triggering test case is derived. To improve the efficiency and effectiveness of SQLancer for bug detection and resolution in various DBMSs, I propose to undertake the following tasks: 1. Implement delta-debugging algorithms for statement reduction in SQLancer to improve efficiency and speed. 2. Define parser-tree transforms to simplify SQL statements. Since SQL syntax varies among DBMSs, generic parser-tree transforms that can be applied across multiple DBMSs will be implemented, such as removing expression nodes or removing certain clauses from the statement. 3. Generalize statement reduction to additional DBMSs by adding bug reproducers to them. The deliverables for this project include: 1. A delta-debugging algorithm that can reduce the bug-inducing statement. 2. Generic parser-tree transforms that simplify SQL statements. 3. Support of statement reduction for more DBMSs.
The project aims to make an electro-mechanical model of a motor which would simulate a motor in gazebo. The purpose is to enhance the motor testing behaviour before real world deployment. It also converts the RC command to motor joint control command more correctly and accuratly.
<p>Spir-fuzz is a C++-based tool that automatically finds bugs in Vulkan drivers. It works by transform the original shader into a new one that is semantically the same. Differences in the output of the new shader and the original one can be caused by bugs in the driver. Thus, our task involves expanding the set of transformations by building sets of transformation classes and writing their corresponding tests and fuzzer passes.</p> <p>Our main task involves the WebGPU Shading Language, a new shading language featured by WebGPU. Since web browsers will have WebGPU, a secure implementation is crucial. To achieve a high test coverage, we use coverage-guided fuzzing. It uses program instrumentation to trace the code coverage reached by each input fed to a fuzz target. The information is then used to make informed decisions that maximize coverage, and thus increase the effectiveness of finding software bugs and security vulnerabilities. This project involves automatic fuzzing using LibFuzzer. Since LibFuzzer-based custom mutators mutate test cases in a domain-specific way, effective designing and implementing Tint-specific custom mutators are essential for this project to succeed.</p>
<p>Panoptic Segmentation aims to further the process of understanding an image by classifying all the pixels in it while also assigning them to unique object instances. I designed and implemented a top-down method by extending the existing MaskRCNN model with a semantic segmentation head to generate both instance and semantic segmentation masks. I also implemented a post-processing layer that combines the resulting masks.</p>
Gammapy currently serializes its data products across multiple data formats, which creates a fragmented workflow with poor metadata handling. This project introduces the Advanced Scientific Data Format (ASDF) for Map and Model objects. ASDF is transparent, human-readable, easily extensible, and supports schemas and validation machinery. This project implements ASDF converters and schemas for Gammapy data objects. The resulting ASDF support provides Gammapy users with a unified, self-describing file format that improves reproducibility and simplifies data sharing across the community.
<p><strong>Productionize API for Wikidata-based Topic Model</strong></p><p>The deadline to create an initial application has passed for the May 2020 Outreachy internship cohort. We are no longer accepting initial applications for internships. We encourage you to sign up for the announcements mailing list to get an email when the next round opens. Why apply to Outreachy? > Start your initial application > May 2020 Outreachy internship cohort > Wikimedia Foundation Community details are hidden until you are approved to participate as a mentor or coordinator.</p><p><br></p><p><strong>Mentorship Cohort:</strong> 2020</p>
<p>Accelerate high-dimensional spatial genomic matrix alignment and multi-sample normalization routines for large-scale spatial transcriptomics datasets.</p><p>Mentee will benchmark memory utilization and implement parallelized matrix transformations using C++ extensions inside BiocParallel.</p><p><br></p><p><strong>Deliverables:</strong></p><ul><li>Parallel C++ kernel extensions for BiocParallel.</li><li>Benchmark suite comparing CPU/RAM throughput.</li></ul><p><br></p><p><strong>Applications closing date: 15-Nov-2025</strong></p>
Musculoskeletal forces from muscles and tendons often wrap around bones and joints along curved surfaces. Current tools rely on numerical approximations for these wrapping paths, which require intensive computation. This project will extend SymPy’s mechanics module by introducing three new wrapping geometry classes: WrappingCone, WrappingEllipsoid, and WrappingToroid each providing symbolic geodesic-length calculations and endpoint tangent vectors. Using closed‑form derivations for cones, series expansions for spheroids, and local unfolding approximations for tori, the new classes will enable users to model complex biomechanical interactions symbolically, offering both exact expressions (where ever available) and good approximations. This extension will empower researchers and educators to explore musculoskeletal pathways analytically within the familiar SymPy ecosystem.
<p>The aim of this project is to integrate an automatic AI-based scheme into the gprMax environment such that it provides real-time FDTD solutions to the user. A Deep Learning-based algorithm can be used to provide real-time solutions based on the user's inputs, thus speeding up the computational process. An AI-based real-time EM solver will be orders of magnitude faster than conventional FDTD and would also alleviate the need for having heavy computational resources at the user’s end.</p> <p>In practice, the goal is to develop user-friendly tools with which the user will be able to effectively parametrise the investigated problem and define the expected range of the parameters. Big data will be generated in an automatic manner to be subsequently used for training a deep learning scheme to predict the electromagnetic (EM) response subject to the parameters of the models.</p>
@nightwatch/mobile-helper is a tool that allows developers and testers to set up a fully functional Android Emulator environment in just a few minutes, without the need to download the complete Android Studio IDE software. This tool comes in very handy, especially for testers who wish to perform end-to-end tests for their website or native app on Android Emulators, by allowing them to set up everything they need in under 5 minutes, as opposed to having to set up the hefty Android Studio IDE and figuring everything out themselves.This project aims to take this tool one step further, by building on its usability and allowing users to do much more with the tool such that it not only benefits the testers but the developer community as a whole, or anyone who wish to quickly set up Android Emulators on their machines, while only downloading what is necessary, saving both internet bandwidth and storage capacity. To achieve this, this project proposes: a) Adding support for wireless adb connection to avoid the inconvenience of using USB cable. Users will be able to connect and manage multiple devices, both real and emulator. b) Adding support for downloading and managing multiple system images and emulators to facilitate use of a wider range of emulator devices. c) Adding functionality to update the installed SDK tools. d) Document all important SDK commands for easy accessibility of the users. The tool will help users to perform all these tasks with a very interactive and easy to use command line interface.
<p>GPS TEC Map (Global Positioning System - Total Electron Count) is an important quantity of the ionosphere for analysis of space weather. Building an accurate predictive model for TEC maps can help in anticipating adverse ionospheric events (ex: solar storms), thereby safeguarding critical communication, energy and navigation infrastructure. I will be building deep spatio-temporal predictive models for GPS TEC Map.</p>
This proposal tackles the challenge of efficiently running machine learning models on resource-constrained HiFi4 DSPs using Zephyr RTOS. The approach begins by deploying and validating open-source ML models (e.g., temperature prediction, micro-speech) on an emulated Xtensa platform before porting and optimizing TensorFlow Lite Micro for HiFi4 DSP using NXP’s AN13970 guide. Deliverables include an optimized ML framework for DSPs, a working sample application, comprehensive documentation, and upstream contributions to Zephyr.
This project aims to develop a more rigorous and focused evaluation testbed than that used by Garrido et al. (2025), with the specific goal of assessing the intuitive physics understanding of Gemma models. By addressing key shortcomings in existing evaluations—such as overreliance on textual outputs, limited control over task difficulty, and incomplete coverage of fundamental physical concepts—this new benchmark will enable more accurate and fair comparisons. Ultimately, it will offer deeper insights into what Gemma models actually grasp about the physical world and guide future improvements in their design and training.
This project aims to enable AI training in a multi-cluster environment while ensuring data privacy through federated learning. In a multi-cluster environment, where different clusters need to access data, federated learning is introduced to safeguard sensitive information. Our goal is to provide a unified interface that allows existing federated learning frameworks, such as OpenFL, NVIDIA FLARE, and Flower, to operate seamlessly in a multi-cluster environment without modifying their existing workflows. These frameworks will continue handling local model training within their respective clusters while aggregating updates into a global model on the hub cluster. To achieve this, we leverage Open Cluster Management (OCM) APIs, including Placement and ManifestWork, to standardize federated learning workflows. By integrating frameworks through a unified interface, we harness OCM’s capabilities to deliver scalable and privacy-preserving AI training solutions in multi-cluster environments.
<p>Sepsis is a potentially life-threatening condition caused by the body's response to an infection. The body normally releases chemicals into the bloodstream to fight an infection. Sepsis occurs when the body's response to these chemicals is out of balance, triggering changes that can damage multiple organ systems. Our main goal here is to train a deep learning model in python using all of its symptoms for the prediction of early onset of sepsis. Depending upon the values fed into the application, a doctor should get a good idea whether a person is susceptible to sepsis and get an early alert which can be critical for diagnosis. The application should be able to make these predictions using only a minimal set of streaming physiological data in real-time. During the course of this project, new deep learning methods, using temporal convolutional neural networks or quasi RNN, a model will be developed to identify markers that predict the onset of sepsis in patients admitted to the intensive care unit. We shall develop this application using the eICU database.</p>
The Amharic DBpedia chapter's past implementation was largely manual, relying on tedious template mapping that has proven impossible to scale given the complexity of the Amharic language and the messy Wikipedia markup. My project's goal is to transition the chapter away from this bottleneck and fully automate the entire Semantic Web pipeline. I'll build an Agentic Orchestration Pipeline (using LangGraph/Mastra) that fixes the problem in three core areas: AI-Powered Mapping: The pipeline will integrate the fine-tuned Afro-XLM-R model to automatically predict and align Amharic properties to the DBpedia Ontology, completely replacing the manual mapping effort. This process is preceded by a Python preprocessor that cleans the raw Wikipedia XML dumps to prevent parser crashes. Human-in-the-Loop Safety (HITL): I will wrap the entire system in a Human-in-the-Loop (HITL) interface, allowing community experts to verify low-confidence predictions to ensure data quality and continuously improve the AI model. Visualization and Accessibility: Finally, I will refactor the current am.dbpedia.org website into a dynamic dashboard that hooks into our deployed SPARQL endpoint, letting users actually query and visualize the live Knowledge Graph. Key deliverables will include the fully functional, automated extraction framework, the publication of the generated Amharic RDF triples to the DBpedia Databus, the deployment of a public-facing SPARQL endpoint (Tentris/Fuseki) for live querying, and the refactoring of the am.dbpedia.org website into a dynamic dashboard for visualizing the new Knowledge Graph.
<p>According to urunc's execution model, the application executes inside a</p><p>VM-based or software-based sandbox. While this model strengthens isolation</p><p>boundaries, it can introduce performance overhead and additional resource</p><p>consumption compared to standard containers. Over the past years, evaluations</p><p>of urunc have focused primarily on spawn time and container density and less</p><p>on other aspects, such as CPU and memory usage, storage overhead I/O</p><p>performance and network latency.</p><p><br></p><p>As a result, it is necessary to conduct a thorough performance evaluation of</p><p>urunc across multiple metrics. The evaluation should span across</p><p>microbenchmarks, macrobenchmarks, and representative real-world workloads.</p><p>Apart from the evaluation itself, it is also important to create a reproducible</p><p>benchmarking suite, with all scripts, tools and documentation, so that anyone</p><p>can extend and reproduce the experiments across different environments and</p><p>versions.</p><p><br></p><p>- Expected Outcome:</p><p> - A comprehensive evaluation report covering startup latency, CPU and memory</p><p> consumption, storage overhead, I/O throughput, and network performance</p><p> across multiple sandboxes (monitor / guest combinations).</p><p> - A reproducible evaluation suite, including all scripts, tools,</p><p> configurations and documentation required to repeat and extend the</p><p> benchmarks.</p><p> - A blogpost summarizing the methodology and the findings of the evaluation.</p><p><br></p>
KubeEdge-Ianvs currently focuses on edge-cloud collaborative learning (training and inference) for a single modality of data. However, edge devices, such as those in autonomous vehicles, often capture multimodal data, including GPS, LIDAR, and Camera data. Single-modal learning can no longer meet the precise inference requirements of edge devices. Therefore, this project aims to integrate mainstream multimodal large model joint learning algorithms into KubeEdge-Ianvs edge-cloud collaborative learning, providing multimodal learning capabilities. Expected Outcome: A benchmark suite for multimodal large language models deployed at the edge using KubeEdge-Ianvs - Modify and adapt the existing edge-cloud data collection interface to meet the requirements of multimodal data collection - Implement a Multimodal Large Language Model (MLLM) benchmark suite based on Ianvs - Reproduce mainstream multimodal joint learning (training and inference) algorithms and integrate them into Ianvs single-task learning - (Advanced) Test the effectiveness of multimodal joint learning in at least one of Ianvs' advanced paradigms (lifelong learning, incremental learning, federated learning, etc.).
<p>Aim is to build all the portions of the FastAI.jl package, inspired by the fastai Python library, which will provide high-level components that can quickly and easily provide state-of-the-art results for tabular tasks, and provide low-level components that can be mixed and matched to build new approaches.</p> <p>This will include handling tabular data of all kinds of format, performing transformations on it if required, creating a model using best practices and entity embeddings, and being able to train the created model.</p> <p>All this will be done without compromising in ease of use, flexibility, or performance, due to the benefits Julia provides, along with the well designed three layered architecture.</p>
Researchers have tremendous difficulty keeping up with the exponential growth of scientific literature, hindering their ability to stay on top of the latest insights in their field. This is especially true in quantum physics, where groundbreaking findings and complex concepts are emerging at a rapid pace, demanding constant attention and analysis from researchers. There is a dire need for AI that can efficiently process and summarize these vast volumes of scientific papers, helping human researchers focus their efforts on the most promising areas of investigation. The problem of knowledge overload and the need for efficient systems to navigate and synthesize vast amounts of literature is not unique to quantum physics. In fact, it is a pervasive challenge across various academic and scientific disciplines, including statistics and data science – the very fields that led to the development of large language models (LLMs) and retrieval-augmented generation (RAG) models. The exponential growth of research papers, coupled with the rapid advancement of statistical methodologies and data-driven approaches, has created a significant knowledge bottleneck in these domains. Researchers and practitioners in statistics and data science often struggle to keep up with the latest developments, hindering their ability to effectively leverage cutting-edge techniques and methodologies in their work. This knowledge overload problem has motivated the development of AI-powered systems like LLMs and RAG models, which aim to streamline the process of knowledge discovery, synthesis, and dissemination. By adapting the proposed "AI Scholar" system to the domains of statistics and data science, researchers in these fields could benefit from personalized recommendations, automated literature analysis, and idea generation capabilities, ultimately accelerating the pace of innovation and fostering cross-disciplinary collaboration.
This project aims to implement a Near-to-Far Field Transformation (NFFT) feature in gprMax. This will enable users to compute the far-field radiation patterns and radar crosssections(RCS) from the near-field data. It will expand gprMax’s capabilities for antenna design, scattering analysis, etc. I propose to implement the Near-to-Far Field Transformation (NFFT) module and its integration at several key points in the existing workflow: - Adding the NFFT command - Module for Field Sampling during FDTD - Post-Processing & Far-Field Calculation - Integration with the API - Validation and Testing - Documentation and Review