Fetching the latest programs, projects, and workspace data.

Research on Multimodal Communication
Showing 5 of 85 projects. Click any project card for scope, mentors, and proposal studio.
Mentors: Student: Mayank Palan
Stentor roeselii is a free-living ciliate species of the genus Stentor. It is a unicellular organism but shows remarkable hierarchy of decisions when faced with a stimulus leading us to question whether there is some level of cognition involved and whether the line between metabolism and cognition is fuzzy implying mitochondria may be the cognitive centre For the last 200 years research into human cognition and decision-making has revolved around rational choice theory, which assumes that actors are utility maximizers capable of searching and finding rational and optimal decisions. However human behavioral data doesn’t support the assumptions that we have infinite time nor infinite cognitive energy to search the full space of candidate choices to produce optimal decisions. Take for example a hiker, who has an infinite number of paths to choose from. It's irrational to assume that the hiker is capable of fully evaluating each of the infinite paths available to them to then choose the most rational/ optimal path. And yet that is the operating assumption for most modern cognitive models. To remedy this, we have the Wayfinding theory which along with an optimal-control based model can be used by researchers in private and academic settings to develop more accurate computational models of cognition in various life species
Mentors: Student: Prakriti Shankar Shetty
Detection of Intonational Units- An overhaul of AuToBI aims to refurbish the AuToBI system first proposed by Andrew Rosenberg in 2009, with today’s state-of-the-art algorithms and models so as to cater to the compute, speed and accuracy requirements that are characteristic to today and the future. The main idea is still the automatic detection and classification of prosodic events (Rosenberg’s thesis considers two events of interest pitch accents and phrase boundaries, but we can target to cater to a wider array of events), but we will lay our focus on the broad goal of establishing a new baseline by employing some of the more sophisticated MLmethods as of today. One of the fairly less-explored avenues of today is the interpretability of models, and with our work, we hope to set precedent for the future of prosody-detection with a highly interpretable model Goals and Deliverables: Goal 1: Try to incorporate speaking rate and speech rhythm as additional prosodic events. Correlated Goal 1A: Analyse the correlation between the prosodic events of speech rate/rhythm with phrase boundaries. Goal 2: Try to incorporate CTC alignment to do away with disparities due to frame-level alignment. Analyse the benefits of the change. Goal 3: Employ techniques like Tandem DNN-HMMs, RNN-Transducers and WSFTs and analyse the benefits, if any Ambitious Goal 4: Employ state-of-the-art techniques to improve speaker diarization. Can try to generalise speakers to larger groups based on ethnicity/region to employ region-specific analysis.
Mentors: Student: Mohamed Ahmed Krichen
This project aims to develop a unified system for sentiment and stance detection in televised news content. Leveraging text, audio, and video modalities, the research seeks to understand the interplay between emotional tone and the position taken on specific issues in news narratives. The goal is to create an integrated tool for comprehensive analysis, contributing to a deeper comprehension of the overall narrative in news broadcasts. The methods include implementing sentiment analysis models for news transcripts, developing stance detection models for specific topics, and building a unified system that combines sentiment and stance information. The project’s potential impact lies in providing researchers and analysts with a valuable tool for nuanced news content analysis.
Mentors: Student: Manish Kumar Thota
This proposal outlines the development of an innovative system designed to enhance video annotation capabilities through the integration of a multimodal vision and language model with spatial-temporal analysis. Utilizing the LLaVA-v1.5-13b model with Video Adapter, known for its rapid inference and entity detection in video frames, this system aims to extract annotations automatically and compile them into a user-friendly CSV format. At its core, a Pydantic API extracts boolean values for annotation entities from enhanced language model responses, facilitating the concurrent processing of multiple videos. The entire workflow, from video and JSON input to CSV output, will be accessible via a Gradio interface on the Hugging Face platform, ensuring ease of use and broad accessibility.
Mentors: Student: Zhongheng Cheng
Despite LLMs' proficiency in various tasks, they struggle with incorporating characteristics like frame blending into sentence generation, a concept well-developed in linguistics and Frame Semantics. The significance of frame semantics lies in understanding words based on conceptual structures, enhancing insights into linguistics, and contributing to fields like economics and politics. The research aims to employ training, fine-tuning, and prompt engineering technologies when using open-source LLMs for generating frame-blending examples. The proposed method involves leveraging FrameNet data for training and fine-tuning LLMs to enhance frame-blending capabilities. The expected outcome is improved frame blending and overall semantic understanding in LLMs, contributing to a deeper exploration of their capabilities and future prospects. Furthermore, a detailed schedule is provided at the end of this proposal.