Fetching the latest programs, projects, and workspace data.
Find open source projects actively accepting contributors. Search repositories, filter by program milestones, difficulty tags, or tech stack.
Use our Orbit AI Matcher to find out! Get instant matching scores based on your developer skills, preferred frameworks, and contribution experience.
Convert your selected open-source project into a winning GSoC, LFX, or Outreachy application using Proposal Studio.
Parents and caregivers increasingly turn to social media for pediatric ENT advice, where misinformation about conditions like ear infections and hearing loss spreads rapidly and without clinical oversight. This project builds a production-ready, open-source AI surveillance system that monitors Reddit, YouTube, TikTok, and Instagram in real time, classifies emerging trends as HARMFUL, CONCERNING, or SAFE using a 6-node LangGraph agentic pipeline grounded in live PubMed evidence retrieval, and surfaces results to clinicians through a Next.js dashboard with SSE alerts and email notifications for high-risk trends. A working 4-node prototype validated on live Reddit and YouTube data is already implemented. The GSoC deliverables extend this foundation with TikTok and Instagram collectors, ASSESS and REPORT nodes, a production Next.js dashboard replacing the existing Streamlit prototype, APScheduler end-to-end integration, Docker Compose deployment, and formal evaluation against a 300-post clinician-labeled reference set targeting precision > 0.85 on HARMFUL and F1 > 0.80 overall.
cBioPortal introduced Similarity Maps, a feature that visualizes UMAP/PCA embeddings of cancer samples, allowing researchers to discover hidden patterns across thousands of samples simultaneously. However, the current implementation is a prototype that fetches data directly from a hardcoded S3 URL, bypassing cBioPortal's data pipeline entirely and following no established datahub standards. This project makes Similarity Maps production-ready by building a proper data layer that follows cBioPortal's existing architecture. This includes designing a datahub file format for embeddings, reusing/extending the existing importer to load embedding data into ClickHouse, and exposing REST API endpoints that replace the hardcoded S3 fetch. The frontend will then be updated to consume the new API dynamically. Additionally, enhanced tooltips and improved colouring options will be implemented to strengthen the visualization. Key deliverables include: an approved database schema, a working EmbeddingImporter.java, a validated datahub file format, REST API endpoints, updated frontend integration, and full documentation
<p>Sugar has been using GTK2 as its primary toolkit, for many years. But recently major changes have occurred in the sugar’s technology which includes an upgrade to a newer version of GTK+ i.e. the GTK3 toolkit and PyGI. Because of these changes in technology sugar is divided on these old and new distributions.</p> <p>This project is all about porting a dozen of major and minor activities of sugar which were based on “GTK2 and GST0”.10 to “GTK3 and GST1.0” in order for sugar to remain innovative stable and fully updated with newer technology.</p> <p>These are the 12 activities whose refactoring will be done : Turtleart, Record, Chat, Calculate, Colors, Stopwatch, Dots and Boxes, Slider-puzzle-branch, Classroom broadcast, Convert, Arithmetic, Pukllanapac.</p>
<p>cBioPortal provides clinical data from a number of different cancer studies. The clinical data from the studies is conveniently organized in a tabular format, however the data available varies between different studies. The variation in data makes it difficult to directly compare all of the different studies. The aim of this project is to standardize the data currently available and to propose certain standards for the data added to cBioPortal from future studies. Standardized data will allow for users to more easily work with the data provided by cBioPortal.</p>
The project aims to design and implement advanced interactive visualizations for ListenBrainz using Nivo for data visualization and integrating with the existing Flask API. Apache Spark will handle efficient data processing and aggregation. These visualizations will offer granular insights into genre trends, artist diversity, and temporal listening patterns, enhancing user experience and engagement. The project will result in the development and integration of the following four interactive charts into ListenBrainz: 1. Generate Artist Listening Activity Statistics: A function to calculate the number of listens for each artist over a specified time range, mapping the data into an interactive visualization that shows the evolution of artists a user listens to over time. 2. Generate Listens by Era Statistics: A function to calculate the number of listens grouped by the year of release of each track, producing a chart that displays listens by era, enabling users to explore the distribution of their music preferences across different periods. 3. Genre-Based Listening Patterns: A dynamic visualization that tracks a user’s engagement with different genres over time, highlighting how their music preferences evolve. It will also analyze trends in genres listened to at different times of day, revealing how mood or activity influences music choices. 4. Top Listeners: A leaderboard-style visualization that tracks and highlights the top listeners sitewide. This chart will showcase the most active users and display who listens the most within their social circles, encouraging community engagement and fostering a sense of connection among users. These interactive charts will enhance the ListenBrainz user experience by providing deeper insights into listening habits and fostering a stronger sense of community.
<p>InterMine has an Android application and is looking to make it available on iOS as well. Native iOS application will allow researchers to have access to InterMine data using fast, reliable and intuitive interface.</p>
<p>Mapping between the natural language text of a contract and attempting to classify data types such as monetary amounts, dates and legal specific terms such as agreement parties, etc. into cicero variables from the existing model library.</p> <p>Therefore, users can export a smart legal contract by natural language contract text.</p> <h3>How to do it?</h3> <h4>A. Data Mapping using NER by RoBERTa.</h4> <p>Build a scaleable NER model by Adapter Transformers based on RoBERTa.</p> <p>The model also have Active Learning pipelines. User can define their own custom data type label then upload data and train the Adapter. By doing so, the model will recognize their new tag.</p> <h4>B. Suggest about templates by Classification Model</h4> <p>When users first upload their natural language contract, NLP model will tell them which smart legal contract template is suitable. So users can use or fine-tune the contract easily.</p> <h4>C. Identification contract variables by BERT QA Model</h4> <p>Input contract text, NLP model will suggest user which they need to put onto smart legal contract's variables.</p> <h4>D. API backend.</h4> <p>User can call the NLP model by API. Provide a Swagger UI Documents and plenty of examples.</p> <h3>For more detail, please <a href="https://github.com/accordproject/labs-cicero-classify" target="_blank">go to README in the GitHub Repo</a></h3>
Problem Statement:- Plone is a robust and secure Content Management System that is used by a wide range of organizations worldwide. However, given the growing amount of information available on Plone websites, it can be challenging to locate the appropriate information fast given the state of Plone. Solution:- The objective of this proposal is to develop a Nuclia Plugin that will scrape websites, extract pertinent data, and index it in the Nuclia Knowledge Box. The Plugin will then be deployed to Plone websites after being merged with the current Plone knowledge Box. Set Of Deliverabilities that will be delivered:- Easy Search without multiple navigations. This will save time and increase productivity. Centralized location for all Plone-related information. Enable developers to explore new techniques, features, and solutions more easily, leading to better results and more efficient development processes. Increase the visibility of Plone and help to attract new users and contributors. Make Plone more user-friendly and increase adoption. Help to build a stronger community by encouraging collaboration.
<p>Teiid is a data virtualization tool. This tool enables querying multiple data sources with a single query. My project covers the following 3 major tasks:</p> <ol> <li>Implementing a translator for parquet files that would enable reading of parquet files from different sources.</li> <li>Implementing data source support for s3 would enable the reading of various files in S3 sources.</li> <li>Implementing data source support for HDFS would enable the reading of various files in the HDFS source.</li> </ol>
The dataset hoster currently returns rows of database rows with no extra information. Over time, the results that are returned from a query have begun to include markup in order to display more information to the user. The data structure which is currently being returned by the hoster should be simplified to make data consumption easier for callers. When an MBID is recognized in a cell of the output, we wish to show a drop-down button next to that MBID. The items of this drop-down should be a link to all of the data-sets that could be initiated with that MBID. For instance, if the user has made a search for a similar artist and the output shows a list of artists, next to each artist MBID, the dropdown should offer to run a new similar artist search and other searches that can be initiated with an artist MBID. These drop-downs would allow users to explore our connected datasets, which would be a great asset. Also, the project has accumulated some bugs over time, many of which remain unaddressed. These bugs need to be identified and fixed as required.
<p>The project has been under constant development for almost a year now. In the past one year, the app has been reworked to compile with the current Android ecosystem. Users can view release information by scanning a barcode, search information about artists, releases, release groups,labels, recordings, instruments, and events , view collections, tag their audio files (to a limited extent) and donate to the MetaBrainz Foundation via PayPal. The application has been released as a beta version on the Play Store for testing purposes. This project aims to make the application stable and robust.</p>
Sugar Labs has more than 250 activities GitHub and elsewhere which have scope for improvement. Since the support for Python2 was withdrawn from Python Foundation, porting the activities to Python3 and GTK+ 3 is very crucial. My work on this project includes, (1) Porting activities to GTK+ 3 and Python3. (2) Implementing basic game design features which attracts the elementary grade children (activities from Honey). (3) Adding activity specific features like collaborations and enhance User Interface & User Experience. (4) Testing and modifying the activities to ensure 12 activities are ready to release.
<p>I will be designing pitch tracker, integrating software keyboard and fixing bugs for music blocks.</p>
<p>Implementation of some of the useful transforms from Vega currently not present in Vega-Lite</p>
The Proposal is how to make updates in the reference implementation based on the harmonization work, the creation of the new FHIR profile, and the revisions to the FHIR Operations. Also examine the GA4GH and FHIR data models, comparing and contrasting different representations of molecular consequences, in order to identify additional harmonization opportunities.
<p>ListenBrainz has recently shifted its statistics infrastructure from Google BigQuery to Apache Spark. Apart from delivering statistics/graphs to end user, using Spark’s cluster computing, its MLlib can be effectively exploited to build an open-source music recommendation system to promote artists across the globe and not only a selected few promoted by the major labels.</p>
<p>Underwater noise pollution impedes the orcas’ ability to communicate and echolocate scarce salmon. To be able to detect orca calls, a deep learning model can be trained, but due to the lack of a vast amount of labeled data, an active learning tool is necessary.<br> This project provides an active learning web application, where the user can visualize the predictions of a deep learning binary classification algorithm on unlabeled data. Users are then allowed to annotate more data by hearing a sample and correcting its label. The project has been built in a modular way to allow the integration with different machine learning models.</p>
<p>The support for Python 2 will be dropped in the coming years and thus it is important to switch to Python 3. There is a need to make sugar-toolkit and activities compatible with Python 3, to ensure the smooth functioning of the sugar activities in the future. The aim of this project is to boost this necessity.</p> <h5>Target</h5> <p>Following are the targets of this project:-</p> <ul> <li>To port all static telepathy bindings to TelepathyGLib.</li> <li>Make activity-chooser window modal.</li> <li>Make sugar-toolkit compatible with both Python 2 and Python 3.</li> <li>Make sugar-desktop compatible with Python 3.</li> <li>Port sugar-activities to Python 3 and make a release.</li> <li>Write the necessary documentation.</li> </ul>
Music Blocks v4 has a modular architecture with a block editor, compiler, and execution engine built separately but not yet connected. This project integrates the Masonry visual block editor into the main application, implements block snapping and connection logic so users can build structured programs, and wires the full execution pipeline (blocks to AST to IR to output) so programs produce real music and graphics. Core deliverables: (1) Masonry module integration with palette and workspace (2) snap-based block connections with visual feedback enforcing programming semantics (3) end-to-end execution connecting blocks to the Painter and Singer modules via a real function registry. Stretch goals include inline block editing and workspace zoom/pan/undo.
Accessing medical data through the care web app can be daunting for patients in rural or remote areas who often face challenges with digital literacy, internet connectivity or lack of access to computers. Currently, retrieving appointments, health records etc requires users to log in and navigate the care web app, which makes it difficult for those with limited digital exposure to access their essential medical information. This project introduces an Instant Messaging wrapper for the care EMR, built as a django plugin. Which allows patients and staff to securely fetch medical records, check appointments etc and receive alerts and notifications through familiar, easy to use messaging apps like WhatsApp. To ensure HIPAA security compliance, the system uses a 2 step authentication process which includes (1). matching phone number and (2). verifying requestor's date of birth. It will also include redis based caching, ratelimiting of requests to prevent abuse and async task management via celery to handle sending of notifications in the background. A frontend plugin will also be developed to give staff the ability to easily send alerts and notifications and allow patients to download PDFs, such as medications and lab reports. Deliverables: 1. A functional, django based backend plugin, capable of securely serving patient and staff queries via various instant messaging providers. 2. A frontend developed using care_hello_fe to manage notifications and handle PDF downloads. 3. Proper tests utilizing pytest and Playwright for both the backend and frontend. 4. Complete and comprehensive documentation using sphinx and swagger, along with setup guides and demo videos. 5. Production ready deployments of both the backend and frontend plugins integrated within care.
<p>The goal of this proposal is to enhance support and usability of <code>cwltool</code> by:</p> <ul> <li>Adding Python 3 compatibility to CWL Tool, Schema Salad and CWL Test projects.</li> <li>Improving mypy annotations and upgrading mypy dependency for CWL Tool and Schema Salad.</li> <li>Reducing <a href="https://github.com/orgs/common-workflow-language/projects/3?fullscreen=true" target="_blank">issue backlog</a> and misc enhancements</li> </ul>
<p>Create Write Activity for Sugarizer : In this project Write activity will be implemented for sugarizer platform which will be in congruence to Write activity on Sugar platform . It is a rich text editor with features like real time multi user Collaboration , Export/import to editable and non editable format , database connectivity , autosave and many more interesting features .</p>
Pippy is the Sugar "learn to program in Python" activity. It comes with lots of examples and has sufficient scaffolding such that a learner could modify an existing Sugar activity or write a new one. The goal of this project is to add "co-pilot"-like assistance to Pippy. A learner should be able to ask the AI to provide example Python code to help them navigate the language and explore possibilities in a more open way than the collection of Pippy examples affords. Besides correcting code and providing example to kids the project aims to help developers navigate through Sugar-toolkit and also GTK basics helping them contribute in more effective way. Our AI-assistant model would provide contributors sufficient examples to understand how the sugar codebase works.