SANE 2026 - Speech and Audio in the Northeast

October 30, 2026

Aerial view of the East Campus of the Massachusetts Institute of Technology (MIT) along the Charles River, 2015, Nick Allen

SANE 2026, a one-day event gathering researchers and students in speech and audio from the Northeast of the American continent, will be held on Friday October 30, 2026 at MIT, in Cambridge, MA.

It is the 13th edition in the SANE series of workshops, which started in 2012 and is held every year alternately in Boston and New York. Since the first edition, the audience has steadily grown to about 200 participants and 50 posters each year, and SANE has established itself as a vibrant, must-attend event for the speech and audio community across the northeast and beyond.

SANE 2026 will feature invited talks by leading researchers from the Northeast as well as from the wider community. It will also feature a lively poster session, open to both students and researchers.

This year, SANE will take place right after the AIR and DCASE workshops, held October 27-29 in Cambridge as well.

Details

  • Date: Friday, October 30, 2026
  • Venue: Thomas Tull Concert Hall, W18 Building, MIT
  • Time: Typically, SANE talks start around 8:30-9am and finish around 5:30-6pm, with an after-party nearby.

Speakers

Click on the talk title to jump to the abstract and bio.

Registration

Registration is free but required. We will only be able to accommodate a limited number of participants, so we encourage those interested in attending this event to register as soon as possible by sending an email to with your name and affiliation.

Poster Session (Submission deadline: 10/10)

If you would like to present a poster, please send an email to with a brief abstract describing what you plan to present. Poster submission deadline is October 10.
Note that SANE is not a peer-reviewed workshop, and you can present work that has been or will be submitted elsewhere (if the other venue does not explicitly forbids it). As SANE participants will be a mix of audio signal processing, speech, and machine learning researchers, the poster session will be a great opportunity to foster discussion as well as get feedback and comments from various perspectives on your most recent work.

Directions

The workshop will be hosted in the Thomas Tull Concert Hall at MIT, located in the new Edward & Joyce Linde Music Building (W18). The address is 201 Amherst St, Cambridge, MA 02142.

 

Organizing Committee

 

Sponsors

SANE remains a free workshop thanks to the generous contributions of the sponsors below.

MERL Google MIT Modulate Moises Bose

Talks

 

Recent Advances in Expressive, Secure, and Health-Centric Speech Intelligence

Berrak Sisman

Johns Hopkins University

In this talk, Berrak Sisman presents recent research from her lab at the Johns Hopkins University Center for Language and Speech Processing (CLSP) on human-centered speech intelligence. The talk highlights their recent progress in controllable speech synthesis and voice conversion, robust speech deepfake detection, mechanistic interpretability of large audio-language models, and clinical applications for speech and language technologies, such as Alzheimer's and depression detection.

Berrak Sisman

Berrak Sisman is an Assistant Professor of Electrical and Computer Engineering at Johns Hopkins University (JHU), where she leads the Smile Lab as part of the Center for Language and Speech Processing (CLSP). She received her Ph.D. from the National University of Singapore. Her research centers on human-centered speech and audio intelligence, focusing on expressive speech synthesis, voice conversion, speech deepfake detection, and the security, privacy, and health information encoded in the human voice. Her research has received multiple recognitions, including the 2026 Speech Communication Best Journal Paper Award, an NSF CAREER Award (2024), Amazon Faculty Awards (2023, 2025), and the Singapore Ministry of Education Tier 2 Award (2023). She is an Associate Editor for IEEE Transactions on Affective Computing, and an elected member of the IEEE Speech and Language Technical Committee (SLTC). She is the Program Co-chair of Interspeech 2026, and she serves on the ISCA Board.

 

Creative Control for Artists in the Age of AI

Bryan Pardo

Northwestern University

The dominant paradigm in generative models for music is to place a huge generative model in a data center and design the interaction to support casual creation with minimal human oversight or input. This emphasis on centralized, low-touch generative modeling for casual users has resulted in a flood of generic music that risks homogenizing our culture and marginalizing the people whose works were used to train the generative models displacing them.
Bryan Pardo’s Interactive Audio Lab chooses a different path: the goal is to support, rather than supplant, human artistry. The tools we develop can be deployed locally, are controllable through multiple modalities (text, audio examples, hand gestures) and support both intuitive control and virtuosity. This gives creators (musicians, sound artists, podcasters, Foley artists) the tools to support new creative work. We also put strong emphasis on attribution and acknowledgment of human artists. In this talk, Bryan will touch on recent work from the lab that supports this ethos and direction.

Bryan Pardo

Bryan Pardo is a professor of Computer Science at Northwestern University. He studies fundamental problems in generative modeling of music, speech and sound effects, computer audition, content-based audio search and also develops inclusive interfaces blind and low-vision users of audio production tools. He is head of Northwestern University’s Interactive Audio Lab and co-director of the Northwestern University Center for HCI+Design. He received a M. Mus. in Jazz Studies in 2001 and a Ph.D. in Computer Science in 2005, both from the University of Michigan. He has authored over 150 peer-reviewed publications. He has developed speech analysis software for the Speech and Hearing department of the Ohio State University, statistical software for SPSS and worked as a machine learning researcher for General Dynamics. His patented technologies have been productized by companies including Bose, Adobe, Lexi, and Ear Machine.

 

On the development of Granite Speech, IBM’s open-source, speech-aware LLMs

George A. Saon

IBM Research

Granite Speech is a collection of open-source, highly accurate speech-aware LLMs specialized in multilingual ASR and bidirectional speech translation for major European languages and Japanese that can be used under a permissive license. They pair Conformer acoustic encoders trained with Connectionist Temporal Classification on publicly available multilingual ASR corpora, Q-former speech modality adapters and Granite text LLMs from 1B to 8B parameters. The role of the adapter is to perform temporal downsampling and to map the acoustic embeddings from the encoder to the embedding space of the LLM. The modality adapter and LLM LoRA parameters are trained jointly with next-token prediction CE loss on several tasks such as ASR, speech translation, ASR with word-level time stamps, speaker-attributed ASR and keyword-biased ASR by varying the text prompt and target transcripts.
The talk will dive deeply into several aspects related to the model architecture, training, and performance of Granite Speech. On the encoder side, we will discuss several optimizations such as block attention, self-conditioning and temporal subsampling within Conformer blocks as well as the use of a novel Muon optimizer during training which are exemplified in our newly released Granite 5.0 TurboCTC encoder-only model. On the speech LLM side, we will discuss balanced sampling, synthetic data generation, speculative decoding and a novel non-autoregressive architecture that edits the CTC hypothesis in a single forward pass through the LLM. Throughout the talk, we will report the performance of our models (accuracy and inverse real-time factor) on the Hugging Face Open ASR leaderboard as well as the Treble Far-Field ASR (FFASR) leaderboard that measures system performance in challenging acoustic conditions such as noise, reverberation and distant microphone speech.

George A. Saon

George A. Saon is Manager of the AI Foundations Speech Technologies group at IBM Research's Thomas J. Watson Research Center in Yorktown Heights, NY, and has been a Distinguished Research Scientist at IBM since 2018. He joined IBM Research in 1998 and has worked on feature extraction, acoustic modeling, speaker adaptation, hypothesis search, and end-to-end sequence transduction for large vocabulary continuous speech recognition. He holds a Ph.D. in Computer Science from Université Henri Poincaré – Nancy 1 (1997) and an engineer diploma from the Polytechnic University of Bucharest (1995). Since 2001, Dr. Saon has been a key member of IBM's speech recognition team which participated in several U.S. government-sponsored evaluations for the EARS, SPINE, GALE, RATS and BOLT programs. He has published over 150 conference and journal papers and holds 20 patents in the field of ASR. Since 2024 he has led Granite Speech, IBM's speech large language modeling work. He is a Fellow of the IEEE and served on the IEEE Speech and Language Technical Committee from 2011 to 2014.

 

Unsupervised Sound Separation Beyond Waveform Regression

Henry Li

Google

Single-channel sound separation typically requires clean, isolated training sources, a luxury unavailable in the wild. While unsupervised methods circumvent this constraint, prior frameworks rely on deterministic waveform regression, leaving them blind to source semantics and prone to regression-to-the-mean artifacts. This talk explores two complementary paradigms that advance unsupervised separation beyond signal-level regression: structured latent representations and continuous generative modeling. First, we show how generalizing mixture-invariant training to semantic embedding spaces (such as CLAP) decomposes mixtures into source-specific latent slots without text supervision, enabling prompt-free separation and fine-grained audio-text retrieval. Second, we derive a self-supervised flow matching framework through the lens of the classic Wake-Sleep algorithm by bootstrapping a generative student flow on teacher estimates. This approach bypasses unconditional prior training while consistently outperforming regression baselines across speech and universal sound benchmarks. Together, these paradigms demonstrate how rethinking representation spaces and generative dynamics enables high-fidelity, prompt-free separation directly from in-the-wild audio. We conclude by outlining failure modes and open challenges: permutation ambiguities along generative trajectories, acoustic mismatches from naive remixing, and the fundamental limits of independent-source priors.

Henry Li

Henry Li is a Research Scientist at Google focusing on generative modeling and acoustic scene decomposition. His work combines theoretical frameworks in inverse problems and flow matching with empirical algorithms to isolate, enhance, and semantically categorize sounds from unlabeled acoustic mixtures. While specialized in audio, his research foundation encompasses generative diffusion models more broadly, including theoretical work on volume-preserving maps, diffusion optimal control, and multimodal generation. His research has been presented at machine learning and vision venues such as NeurIPS, ICLR, ICML, and CVPR. He earned his Ph.D. in Applied Mathematics and a B.S. in Mathematics and Computer Science from Yale University. When not untangling audio mixtures, he enjoys running, drumming, and hiking mountains of decidedly moderate height.

 

Building Worlds of Sight and Sound

Ruohan Gao

University of Maryland, College Park

Our everyday world is inherently multisensory: what we see shapes what we hear, while what we hear reveals physical properties and spatial structures that vision alone cannot capture. In this talk, I will explore how we can build computational worlds that jointly model sight and sound, progressing from individual objects to static spaces and, ultimately, dynamic environments. I will first discuss audio-visual objects, including differentiable impact sound modeling and our recent work that leverages visual information to model and render sounds from physical interactions. I will then move from objects to spaces, showing how visual cues from multi-view images can inform spatial acoustic modeling and enable the reconstruction of acoustic environments. Finally, I will discuss our recent efforts toward 3D and 4D audio-visual scene generation, including methods that jointly model visual appearance, geometry, static and dynamic sound sources, and spatial sound fields to create explorable environments that can be both seen and heard. Together, these efforts point toward machines capable of understanding, reconstructing, and generating rich worlds of sight and sound.

Ruohan Gao

Ruohan Gao is an Assistant Professor of Computer Science at the University of Maryland, College Park, where he leads the UMD Multisensory Machine Intelligence Lab. Before joining UMD, he spent a year at Meta as a Research Scientist after his Postdoc at the Stanford Vision and Learning Lab. He received his Ph.D. in Computer Science from The University of Texas at Austin. He mainly works in the fields of computer vision and machine learning with a particular emphasis on multisensory machine intelligence involving sight, sound, and touch. His research has been recognized with the Sony Faculty Innovation Award, AAAI New Faculty Highlights, the Michael H. Granof Award for UT Austin’s Top 1 Doctoral Dissertation, the Google PhD Fellowship, the Adobe Research Fellowship, a Best Student Paper Award at WASPAA 2025, a Best Paper Award Runner Up at BMVC 2021, and a Best Paper Award Finalist at CVPR 2019.

 

Towards Fully Open General Audio Intelligence

Sreyan Ghosh

Google

What does it take for AI to understand the world through audio? In this talk, I will present our work on fully open large audio-language models, highlighting advances in architectures, training curricula, and internet-scale data curation. I will explore how unified encoders connect speech, sound, and music, and how new data and training infrastructure extend audio understanding from short clips to 30-minute recordings. Along the way, expert-level benchmarks reveal persistent gaps between current models and human audio reasoning. Finally, I will show how audio enriches real-world video comprehension, and discuss our work integrating audio into world models. Together, these directions highlight why listening is essential to AI systems that understand and model the world.

Sreyan Ghosh

Sreyan Ghosh is a Research Scientist at Google DeepMind, where he works on audio and audio-visual intelligence for Gemini. His research focuses on speech-to-speech modeling, speech translation, large audio-language models, audio representation learning, and long-form audio understanding and reasoning. He completed his Ph.D. in Computer Science at the University of Maryland, College Park, advised by Dinesh Manocha and Ramani Duraiswami. During his Ph.D., he co-led the development of the Audio Flamingo family of models and the MMAU series of audio understanding and reasoning benchmarks. Previously, he held research internships at Nvidia, Adobe and Microsoft. He is a recipient of the NVIDIA Graduate Fellowship and UMD’s Outstanding Graduate Assistant Award.