Fetching the latest programs, projects, and workspace data.
Find open source projects actively accepting contributors. Search repositories, filter by program milestones, difficulty tags, or tech stack.
Use our Orbit AI Matcher to find out! Get instant matching scores based on your developer skills, preferred frameworks, and contribution experience.
Convert your selected open-source project into a winning GSoC, LFX, or Outreachy application using Proposal Studio.
<p>The main goal of the project is to reproduce selected key findings from the empirical asset pricing literature and related investment practices. Among our background references there are Ilmanen (2011) and Engle et al. (2016), both focused on central issues in the subject. The project pertains to the <em>Quantitative Analysis of Active Portfolio Management</em> and the three broad areas under which implementations fall are: (i) approaches to dynamic asset weighting, (ii) return factors and their risk premia and (iii) time-varying expected returns. An integral part of the project consists in collecting financial data of interest, including Profs. Fama-French's academic library, replication data sets underlying research produced at companies such as AQR or Ibbotson Associates (Morningstar) and data from providers like Bloomberg, CRSP, and MSCI Barra. Real-world quantitative data analyses follow both to reproduce target papers and show functionality. Our work is bundled into an open source <code>R</code> package, released on GitHub under the AGPL. We believe that providing a modeling and validation toolbox will add value for the users of the field by allowing them to freely produce and reproduce research.</p>
<p>Semantic saliency detection be implemented based on CNN module. Quantization method described in <a href="https://petewarden.com/2016/05/03/how-to-quantize-neural-networks-with-tensorflow/" target="_blank">Pete's blog for quantization</a> will be added into project <a href="https://github.com/nyanp/tiny-cnn" target="_blank">tiny-cnn</a> together with deconvolutional functions. This project will be the dependency of OpenCV afterwards. There is also an IOS APP demo for such tiny deep learning structure.</p>
<p>Meshery Models are declarative representations of infrastructure, applications, and their relationships - the canonical artifacts through which Meshery understands and manages cloud native systems. Today, Meshery lacks a standardized, portable distribution mechanism for these models. OCI registries (Docker Hub, AWS ECR, GitHub GHCR, and others) have emerged as the universal artifact store for the cloud native ecosystem, and [ORAS](https://oras.land) (OCI Registry As Storage) provides the Go-native tooling to push and pull arbitrary artifacts to any OCI-compliant registry. This internship implements end-to-end OCI registry support for Meshery Models - from new Connection and Credential types for major registries, to ORAS-powered push/pull logic in the Meshery server, to a redesigned Registry page in the Meshery UI that gives users full visibility and control over their model artifacts across registries.</p><p><br></p><p>Recommended Skills: Golang, REST API development, React. Familiarity with OCI image specifications, container registries, or ORAS is a plus. Experience with Meshery or other CNCF projects is welcomed but not required.</p><p><br></p><p>Responsibilities:</p><p> - Design and implement Connection and Credential types for Docker Hub, AWS ECR, GitHub GHCR, and additional OCI-compliant registries within Meshery's existing connection framework.</p><p> - Implement Golang server-side logic using the ORAS SDK to push and pull Meshery Models (and their component schemas, relationships, and policies) to and from any OCI-compliant registry.</p><p> - Define the OCI artifact media types, manifest structure, and layer conventions used to package Meshery Models for registry storage.</p><p> - Enhance or rewrite the Registry page in Meshery UI to surface connected registries, browsable model artifacts, push/pull controls, and credential management.</p><p> - Write integration tests covering push, pull, and round-trip fidelity of Meshery Models across at least two registry backends.</p><p> - Document the new registry integration, artifact format, and UI workflows in Meshery's official documentation.</p><p><br></p><p>Expected Outcome:</p><p> - Meshery users can connect to Docker Hub, AWS ECR, GHCR, and other OCI registries using managed credentials and push or pull Meshery Models directly from the Meshery UI and `mesheryctl`.</p><p> - A well-defined OCI artifact convention for Meshery Models, documented and suitable for adoption by the broader Meshery ecosystem.</p><p> - A redesigned Registry UI page providing a unified, registry-agnostic interface for model artifact management.</p><p><br></p>
This proposal aims to add a variety of scientific utilities to the pvlib python project that will focus on new models and research tools regarding shading and spectral effects. These improvements will contribute to expanding the current capabilities in simulating, analyzing, and researching photovoltaic systems. The proposed improvements are comprised of a shading losses model with an example of row to row trackers shading example; a variety of utilities for researching and estimating spectral mismatch losses; a non-uniform irradiance mismatch losses model, and models for Photosynthetically Active Radiation (PAR). These last models would be used in agrivoltaics, the integration of photovoltaic systems in crops. Further details and references can be found in the Proposal PDF file.
Large Language Models are picking pace very quickly and they are turning out to be extremely good in multiple tasks. With the help of zero-shot, few-shot, and fine tuning techniques we could effectively specialize a language model for the use case. Summarization is one such use case that has been widely researched for a couple of years now. Broadly there are techniques such as Abstractive and Extractive approaches. The motive of this project proposal is to handle the summarization task (mostly Abstractive + Extractive hybrid approach) through the language model’s (foundation model) lens. This project aims to cover everything from data collection, EDA, experimenting with different language models to developing production-scale system that can take GitHub repo as reference and provide summary. One of the challenges that is novel is to use smaller sized models to achieve great performance in summarization. SCoRe Lab has been into developing solutions in the space of making user life easier with products such as D4D, Bassa, Track Pal, and others. This project will add to that portfolio and would be a great reference for AI practitioners and system developers which aims to work right from data to production-grade end product using AI and Systems.
Currently, ns-3 provides the TR 38.901 channel model as the default option for sampling MIMO wireless channels which take into account the presence of (possibly directional) multiple antenna elements at both transmitter and receiver. However, this model represents a major performance bottleneck for most end-to-end simulations, especially whenever they involve upwards of hundreds of nodes. Accordingly, the goal of my proposal is twofold: (i) to provide ns-3 with a simplified statistical channel model, to be used for analyses which are focused on the upper layers of the protocol stack only ; and (ii) to improve the performance of the TR 38.901-based framework which is currently available in ns-3, for the remainder of the use cases.
<p>CTLearn is a Python package for using deep learning to perform analysis tasks on data from imaging atmospheric Cherenkov telescopes (IACTs). These tasks may be either classification or regression problems. The sensitivity of IACTs is mostly driven by our capability of distinguishing gamma-ray induced events from cosmic-ray induced events. The better our deep learning models are in telling these two populations apart, the better our reach in the gamma-ray Universe will be. Therefore, optimizing our deep learning models has the potential to make a difference in our view of the Universe at these energies. This project aims to implement an automated model optimization framework in CTLearn, including random and grid search, Bayesian and genetic algorithms based optimization.</p>
Classical diffusion models (DMs) have experienced rapid growth in usability, availability, and research. The classical DM algorithm consists of two main parts: a forward diffusion process and a learned denoising step. The forward diffusion process consists of discretely applying noise to an image until it is unrecognizable. The denoising process reverses the noise, predicting the original image. The latter is the biggest bottleneck in model training and is likely to benefit from the implementation of parametric quantum circuits (PQCs), creating a hybrid quantum DM. This model would be trained on multiple high energy physics (HEP) datasets, such as the Quark-Gluon dataset for generating quark and gluon jets, alongside a classical DM, and the results could be evaluated taking into account resources, time, and quality.
Fast and accurate simulation of high-energy physics experiments is crucial for advancing our understanding of fundamental particles and forces in nature. However, the traditional method of obtaining such simulations using Monte Carlo methods is computationally expensive. Recently, machine learning techniques such as generative modeling have been proposed as an alternative approach to providing simulation results. In particular, diffusion models have emerged as highly accurate, but the standard formulation of the diffusion process suffers from slow inference speeds. This project aims to accelerate the inference of diffusion models by exploring recently proposed techniques in the field, such as distillation, continuous-time diffusion formulations, or more efficient architectures. The methods developed in this project will be integrated into the Geant4FasSim repository, enhancing its capabilities for fast and accurate simulation of high-energy physics experiments.
The project I am proposing would expand DeepChem’s tools for working with genomic datasets for drug discovery thus strengthening DeepChem’s new Bioinformatics initiatives. I will be implementing a state-of-the-art predictive model for regulatory genomics and adding the relevant datasets for testing. As part of my project, I would compose a tutorial overview on interpreting regulatory sequence data using deep learning. I will figure out what loaders and featurizers to use to translate genomics data into numerical representations that machine learning models can understand. I will also implement gkm-SVM so that it is easier to develop other models down the road that have this dependency. A big part of my project will be identifying how to leverage DeepChem's infrastructure towards biomedical questions informed by genomics as well as identifying future areas for development.
Physicists at the LHC need to perform ML inference on massive amounts of data. Currently, the bookkeeping of ML model files is an unsolved problem. The goal of this project is to evaluate the CernVM File System (CVMFS) as a platform to store, organize, and distribute model files. The two primary issues to tackle are latency and infrastructure. We know there is a latency with using CVMFS versus local storage, so I will rigorously benchmark the performance of CVMFS and determine ways to minimize this overhead. For infrastructure, I will test integrating different services with CVMFS such as Cern Document Server (CDS) or Kubeflow's KServe. The final deliverables will include extensive documentation on benchmarks and best-practices for using CVMFS, and an overall evaluation of CVMFS and the possible integrations like KServe. An example deployment of ML models from the ML4EP project using CVMFS will also be demonstrated.
<p>We create an R Package to run full Bayesian inference on Hidden Markov Models (HMM) using the probabilistic programming language Stan. By providing an intuitive, expressive yet flexible input interface, we enable non-technical users to carry out research using the Bayesian workflow. We provide the user with an expressive interface to mix and match a wide array of options for the observation and latent models, including ample choices of densities, priors, and link functions whenever covariates are present. The software enables users to fit HMM with time-homogeneous transitions as well as time-varying transition probabilities. Priors can be set for every model parameter. Implemented inference algorithms include forward (filtering), forward-backwards (smoothing), Viterbi (most likely hidden path), prior predictive sampling, and posterior predictive sampling. Graphs, tables and other convenience methods for convergence diagnosis, goodness of fit, and data analysis are provided.</p>
The current privilege model of the daemon SWHKD is a bit rough and has issues concerning the security and the UX ascepts related with it. The current problem revolves around gaining access to hardware safely and efficiently, essentially to replace the current privilege model with a simpler and robust system. To do this, I plan on replacing the current IPC oriented model with an ITC (Inter Thread Communication) based, main / child thread model that is known to be more secure and would vastly improve SWHKD
<p>Competing Destination (CD) models, which are an extension of Spatial Interaction (SI) models, has been around since 1980's and are often used within Economic and Social Sciences . CD models involve the analysis of flows from an origin to a destination similarly to traditional SI models, however, CD includes the ‘Competing Destination’ term also known as ‘Accessibility’ term, which accounts for the spatial-structure effect of the flows from behavioural perspective. Although the SI models are established within Python and R modules already (SpInt, simR), specification for deriving Accessibility term for CD estimation is missing. This project aims to fill this gap by developing a Competing Destination class that will include the computation of the accessibility term and will be binded to the existing SpInt module. The main challenge here is the scaling of the accessibility computation for large datasets. This is an important aspect of the project as flow datasets often have hundred thousands and even millions records. This project will extend the use of the SpInt module and PySAL library for high level analysis of spatial flows.</p>
Accurate simulation of particle interactions within detectors is one of the most computationally expensive tasks in High Energy Physics (HEP). While classical Monte Carlo simulators like Geant4 are highly accurate, they are too slow to meet the unprecedented data generation demands of the upcoming High-Luminosity LHC. Quantum Generative Models, specifically Quantum Diffusion Models, have been proposed as a faster alternative. However, current pure-quantum approaches struggle to scale beyond low-resolution "toy" datasets due to qubit limitations, deep circuit noise, and training instabilities. This project proposes the development of a hybrid Quantum Latent Diffusion Model (QLDM). By leveraging a classical Variational Autoencoder (VAE) to compress high-dimensional calorimeter data into a lower-dimensional latent space, the quantum diffusion process can be executed efficiently without overwhelming quantum resources. The final deliverable will be an open-source pipeline capable of generating high-resolution physics simulations. This approach allows the quantum model to operate in a regime where its ability to represent complex probability distributions may provide an advantage over classical generative models.
<p>DBNQA is a large database that can be used to create question answering models. Using a question answering model like NSPM, we can create and experiment with various kinds of QA models with DBNQA. But as we create more and more models it’s become more complicated to manage and use these models. To overcome a situation like this, it’s mandatory to maintain an AI lifecycle management mechanism. This project is trying to address this problem by designing and implementing a lifecycle management framework for DBpedia Neural Question Answering Models.</p>
The current validator of component model inside of WasmEdge only check nested module and ensure VM can run the nested modules without problem, but the validations from component model are mostly skipped. Expected Outcome: 1. One should create a workable (merged into upstream) implementation of validator by working on 2. include/validator/validator_component.h 3. lib/validator/validator_component.cpp 4. The visitor pattern are already setup. Recommended Skills: Since component model proposal separate their validation spec, one should able to find requirements from https://github.com/WebAssembly/component-model/tree/main/design/mvp
Enhance Meshery's existing orchestration capabilities to include support for kro ResourceGraphDefinitions (RGDs) as first-class Meshery Models. This involves enabling Meshery to manage and orchestrate RGDs, similar to how it handles other Kubernetes resources. The project will also include generating support for ResourceGraphDefinition in Meshery's Model generator. Expected Outcome: - Meshery will be able to orchestrate and manage kro RGDs. This includes the ability to deploy, configure, and manage the lifecycle of RGDs through Meshery. The Meshery Model generator will be updated to automatically generate models for kro RGDs, simplifying their integration and management within Meshery. This will be an officially supported feature of Meshery.
<p>The goal of this project is to develop a facility location modeling module that supports various distance measures and returns an optimal solution to the problem. The <a href="https://pysal.org/spaghetti/notebooks/facility-location.html" target="_blank">example</a> provided will serve as a base for this project. This module proposed will support four models:</p> <ul> <li>Location Set Coverage Model</li> <li>Maximal Set Covering Model</li> <li>p-Median Model</li> <li>p-Center Model</li> </ul>
There is a spurt of pretrained edge machine learning models that can be quickly deployed in microcontrollers, thanks to TensorFlow Lite. Reducing RAM usage and increasing battery life are a critical aspect in edge ML models. The core idea of our proposal is to prepare Tensorflow Lite libraries/examples that would reduce energy consumption based on machine learning architecture optimisations. As an example, we created a video that demonstrates how simple machine learning and microprocessor peripheral optimisation can reduce the current consumption of the Wake Word detection model in the Arduino Nano 33 BLE Sense microcontroller. These optimisations show that the Wake word detection model can run for upto 12 days with a 240 mAh coin cell battery.
This project aims to enhance PVNet, a multi-modal deep learning model for solar energy forecasting, by integrating aerosol data as an additional input feature. Aerosols like dust, smoke, and haze significantly affect solar irradiance but are not currently accounted for in PVNet’s forecasting pipeline. The work involves identifying suitable open aerosol datasets, preprocessing and aligning them with existing NWP inputs, and incorporating them into the model architecture through a dedicated encoder. By comparing baseline and aerosol-aware models, we aim to assess the impact of aerosols on forecast accuracy. The final deliverables will include updated data pipelines, modified model components, and performance evaluation. This project contributes toward more accurate, climate-aware solar forecasting using open-source tools and public datasets.
<p>Multi model databases are becoming more popular due its native support for polyglot persistence. Polyglot Persistence is a term to mean that when storing data, it is best to use multiple physical data models, chosen based upon the way data is being used by individual applications or components of a single application. Multi model databases acknowledge the need for multiple data models, combining them to reduce operational complexity, operational costs, extensibility and maintain data consistency. Apache Gora currently supports OrientDB datastore as a multi model database, the project proposes further extending multi model database support with ArangoDB</p>
<p>The main objective of this project will be to implement various neural network models and to improve the documentation for the TensorFlow User community. This will involve both re-implementation of existing models in tf2.x and also implementation of new models.</p>
To implement some of machine/deep learning models on Minerva OSS Dataset from FOSSology and integrate it to atarashi as an agent. Models that are going to be used are logistic regression, LSVM, Naive baiyes, doc2vec for semantic similarity and bert model using fine-tuning.