Fetching the latest programs, projects, and workspace data.
Find open source projects actively accepting contributors. Search repositories, filter by program milestones, difficulty tags, or tech stack.
Use our Orbit AI Matcher to find out! Get instant matching scores based on your developer skills, preferred frameworks, and contribution experience.
Convert your selected open-source project into a winning GSoC, LFX, or Outreachy application using Proposal Studio.
The goal of this proposal is to address the issue of how difficult it can be to package and reuse computational workflows and analyses in genomics and health research since diverse data and workflow standards don't always work together seamlessly. By creating a Python library and related command-line tool that enable bidirectional conversion between RO-Crates (with pertinent computational workload profiles) and GA4GH WES/TES payloads, the RO-GA4GH Bridge project seeks to address this issue.
The cBioPortal provides access to a wealth of cancer genomics data, but new users often struggle to navigate its features, and existing users may not be aware of all available tools. This project aims to enhance the user experience of the cBioPortal for Cancer Genomics by implementing an interactive web tour feature. The objective is to guide new users on how to use the portal and to showcase new features to experienced users. The project proposes using reactour for implementing the web tour and local storage to record the tour progress. The deliverables include implementing the web tour feature for the group comparison tool, allowing users to disable hints, and addressing difficulties such as automatically filling input boxes and continuing tour steps after jumping to other pages.
<p>The Canadian Common CV (CCV) is a tool that allows Canadian researchers to input their resume in a standardized format. It is used by multiple organizations such as granting agencies, federal, provincial and academic institutions. The tool enables users to output results in an XML document, that can be used in external applications.</p> <p>The main objective is to build an application that can ingest, store and query data from the CCV and show the results in a dashboard. The Web Application will be developed in line with a well documented, tested feature list, able to work across on all browsers.</p>
<p>The problem of population genetics can be viewed as a stochastic process. The aim of the project is to understand the reverse transition dynamics of the system conditioned on the end position. In order to understand the feasibility of using reinforcement learning to the problem, agents are tested against stochastic processes that resemble population genetics. Practical evidence is validated to provide a sanity check on the feasibility of deployment of such a method in practice in large-scale problems of population genetics.</p>
<p>StarFix is a cross-platform client-side application that would let you open a file in the Editor of your choice (vscode, eclipse, intellij, emacs, vi, etc.) and other commands locally directly from the browser or file system. This will enable an option to “Open in IDE” through browser extension on various websites like Github, Gitlab, etc similar to the one you see in this demo below. This season we aim to improve the existing tool by resolving the existing issues and adding new features that the user may need .</p>
<p>A key step towards the development of effective personalized cancer treatments is the collection of vast stores of patient data on a variety of tumor types, and cBioPortal has taken these efforts by The Cancer Genome Atlas (TCGA) and provided an accessible visualization tool for biological data discovery. Currently, cBioPortal has a multitude of information on DNA and RNA, but their information on protein levels derives from protein arrays, a technology that has severe limitations relative to mass spectrometry. Clinical Proteomic Tumor Analysis Consortium (CPTAC) has performed mass spectrometry on the patient samples for TCGA tumors and it also includes information on post-translational modifications. I propose to integrate this information with the cBioPortal visualization interface. This involves extending their database with protein abundances and PTMs based on the CPTAC data and updating the cBioPortal interface with a new tab for post-translational modifications and updating existing tabs to accommodate the new protein data. By providing this upgrade, a tool that helps both scientists and doctors see personal and population biology at a glance will gain a new dimension.</p>
Currently, User have to face complex workflow for simple use cases (such as “adding a book”) with lots of repetition of data. This project aims to design and implement simple yet powerful editor for such cases which will abstract away bookbrainz specifics from a user by introducing a new “book” template/interface which will wrap over existing edition entity with useful defaults and pre-configured relationships. This will also allow experienced user to create multiple entities and link them in same workflow.
<p>On MusicBrainz, URLs are external links attached to entities; relationships define their role and include metadata like date period, they provide extra information about the entity and inspire users to explore further.</p> <p>Currently, the URL relationship editor on MusicBrainz has many shortcomings. Though it is a relationship, the URL editor has a different UI from other editors. Automatic rules were built to ease the effort, but outdated and buggy rules make them an obstacle in some situations. For better user experience, we would like to enhance the URL relationship editor, to support multiple relationship types for a single link, to provide a unified appearance, better user experience and more flexibility.</p>
Plone’s pas.plugins.authomatic add-on integrates various social authentication providers into Plone using the Authomatic library. Many providers, such as Google, Facebook, Twitter/X and GitHub, LinkedIn and Amazon have since updated their APIs, requiring certain revisions for secure and compatible login integration. This project aims to: • Update the plugin to support the latest APIs of major providers. • Add support for modern OAuth2 flows and remove deprecated methods. • Improve test coverage and documentation for future maintainability.
The project idea aims to reduce power consumption on Flotta-device agents on small factor edge devices at several levels. OS toolings would be implemented to get data about the energy consumption at the CPU level of the device agent using a Custom Kepler monitoring system and a power meter. After the confirmation of the data, these parameters would be passed to the control plane via metrics using Prometheus as Internal Prometheus TSDB is already used by the flotta-device agent, Required research on variation in energy consumption and performance with the number of workloads and resources allocated to them would be done and the obtained findings is to be summarised in a blogpost. The research part would be helpful for the development of better ideas for workload allocation and containers for Flotta, and IoT based container workload projects in general. Two energy profiles would be developed namely Flotta-PowerSaving Mode and Flotta-UltraPowerSaving Mode which would aim to turn - off specific kernel modules and operations which are not feasible and are unnecessary for running container workloads at low energy, The data from the research would be integrated with these for determining the ranges in which it would operate.
MusicBrainz contains data about upcoming and past events. This data is not surfaced, nor available to ListenBrainz users in any way. This data should be utilized in ListenBrainz to provide it's users with an easy way to discover events. This will be achieved by copying over, and caching the event data from MusicBrainz in a suitable format, and creating endpoints, as well as adequate frontend to display this information. The user should be able to discover events, gather event data from ListenBrainz itself, and find events relevant to them based on their preferred artists.
<p>Single-Cell sequencing is a new methodology that provides high-resolution data to understand individual cells. Applied to RNA sequencing, it can reveal a multitude of relevant information, such as identifying different cell types within a sample and RNA velocity to understand cell fate. The possibility to query individual cells permits a fine-grained analysis compared to conventional NGS (Next Generation Sequencing) methods.</p> <p>Reactome is a “free, online, and open-source database” that stores pathway data from various organisms, such as genes that make up the pathways and hierarchy mappings. Leveraging on this curated data and using single-cell resources from descartes database, we can enhance this data by annotating which pathways each identified cluster represents.</p> <p>This project aims to create a dockerized python pipeline to extract pathway activity from single-cell clusters in a systematic manner, annotating each cluster with which pathway it represents.</p>
<p>This project proposal is based on the “Improve the Customization of Plone Listings” -idea listed on project ideas gathered by the Plone community. As stated in the suggestion, Plone Mosaic add-on currently lacks a way to display customized listings on pages. Customizable listings and search pages are stated to be two often requested features. My proposal is to start solving this problem by allowing users to customize the content of tiles with same WYSIWYG-functionality as already used in mosaic.</p>
<p>Underwater noise pollution impedes the orcas’ ability to communicate and echolocate scarce salmon. To be able to detect orca calls, a deep learning model can be trained, but due to the lack of a vast amount of labeled data, an active learning tool is necessary.<br> This project provides an active learning web application, where the user can visualize the predictions of a deep learning binary classification algorithm on unlabeled data. Users are then allowed to annotate more data by hearing a sample and correcting its label. The project has been built in a modular way to allow the integration with different machine learning models.</p>
<p>The VCF (Variant Call Format) is a format for text files, which is generally stored in a compressed manner to make the data retrieval of variants fast. The data which is redundant is not stored, only the variations are stored. VCF files are used to store all variant types which includes single nucleotide polymorphism (SNP) in a specific position of the genome, short insertions and deletions (INDEL) and structural variants (SV). VCF-validator includes various checks to ensure that the VCF file is consistent. It is based on a formal grammar and performs lexical, syntactic, and semantic analysis of the VCF file. It also includes a tool called VCF-debugulator which fixes errors such as the presence of duplicate variants automatically. SNPs and INDELs are fully supported in VCF-validator, but the support for SVs is still limited. The aim of this project is to improve the support for structural variants in the validator and the debugulator.</p>
This project aims to create a new, user-friendly interface for building and managing workflows in Plone. Workflows in Plone manage content sharing and editing by setting statuses (private, pending review, published) and controlling access permissions. The final outcome will be a Plone add-on, complete with thorough tests and documentation, allowing users to to create secure, efficient workflows directly in Volto.
<p>I would like to work on issue #70, “OncoKB Analysis in Study View”. An inconvenience for both biologists trying to get valuable information to make informed decisions on prognosis and patients who would like to readily have access to this information for their own comfort, is the time needed to access this data. As the genomic data sets for each study vary in size, viewing annotations for studies that have large data sets in real-time is a challenge because it takes a while for the information to load. For example: MSK-IMPACT 2017 has a data set of 10,000 samples, so it will take a few minutes for the study view to load. I would like to work on the three backend tasks and first frontend task. This means pre-annotating the mutation data and storing it into the MySQL database and creating an additional column in the mutation table called “annotation”. Querying and using the web API to provide the information needed to solve the front-end goal of creating pie charts for oncogenicity and highest level of sensitive therapeutic implications per sample.</p>
The current logging system for Matrix communications at MeB relies on an unmaintained IRC-based bot (BrainzBot) that logs messages to a PostgreSQL database and displays them on chatlogs.metabrainz.org. While this system works, it uses outdated technology with known vulnerabilities, doesn't support many of Matrix's features and is no longer suitable as Metabrainz' primary communication platform. This project proposal replaces BrainzBot with a new archival service that archives messages directly from Matrix to HTML files and a PostgreSQL database. It will support Matrix features like message editing, reactions and media, and provide full text search over all messages. Both historical and new messages as they come in will be archived.
This proposal aims to create a client library in rust and command line interface (CLI) for interacting with GA4GH environments in python. It will be easy to upgrade and maintain different versions of API’s with the CLI, which will be done by generating code using openapi-generator automatically. It also involves cross language bindings to access the functions in other languages like C, Python, Go and JavaScript. It will make it easier for organisations and users to access GA4GH standards, increasing the usability of GA4GH standards worldwide
The project aims to design and implement advanced interactive visualizations for ListenBrainz using Nivo for data visualization and integrating with the existing Flask API. Apache Spark will handle efficient data processing and aggregation. These visualizations will offer granular insights into genre trends, artist diversity, and temporal listening patterns, enhancing user experience and engagement. The project will result in the development and integration of the following four interactive charts into ListenBrainz: 1. Generate Artist Listening Activity Statistics: A function to calculate the number of listens for each artist over a specified time range, mapping the data into an interactive visualization that shows the evolution of artists a user listens to over time. 2. Generate Listens by Era Statistics: A function to calculate the number of listens grouped by the year of release of each track, producing a chart that displays listens by era, enabling users to explore the distribution of their music preferences across different periods. 3. Genre-Based Listening Patterns: A dynamic visualization that tracks a user’s engagement with different genres over time, highlighting how their music preferences evolve. It will also analyze trends in genres listened to at different times of day, revealing how mood or activity influences music choices. 4. Top Listeners: A leaderboard-style visualization that tracks and highlights the top listeners sitewide. This chart will showcase the most active users and display who listens the most within their social circles, encouraging community engagement and fostering a sense of connection among users. These interactive charts will enhance the ListenBrainz user experience by providing deeper insights into listening habits and fostering a stronger sense of community.
<p>InterMine has an Android application and is looking to make it available on iOS as well. Native iOS application will allow researchers to have access to InterMine data using fast, reliable and intuitive interface.</p>
<p>Teiid is a data virtualization tool. This tool enables querying multiple data sources with a single query. My project covers the following 3 major tasks:</p> <ol> <li>Implementing a translator for parquet files that would enable reading of parquet files from different sources.</li> <li>Implementing data source support for s3 would enable the reading of various files in S3 sources.</li> <li>Implementing data source support for HDFS would enable the reading of various files in the HDFS source.</li> </ol>
Extend the "Chart" functionality on the Study View Page by supporting additional types of charts such as line charts and area charts. Support chart toggle functionality to allow users to select different types of charts for the particular data set. Implement "group chart" functionality to improve UX for chart analysis.