BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//EVENY//SV//
CALSCALE:GREGORIAN
METHOD:PUBLISH
BEGIN:VEVENT
UID:aug-6-audio-and-ai-meetup-stockholm-20260806@eveny
DTSTAMP:20260729T034600
DTSTART;TZID=Europe/Stockholm:20260806T180000
DTEND;TZID=Europe/Stockholm:20260806T200000
SUMMARY:Aug 6 - Audio and AI Meetup
DESCRIPTION:Join our virtual meetup to hear talks from experts on cutting-edge topics across AI\, ML\, and computer vision. **Date\, Time and Location** Aug 06\, 2026 9:00 AM - 11:00 AM PST **[Online. Register for t…\n\nJoin our virtual meetup to hear talks from experts on cutting-edge topics across AI\, ML\, and computer vision.\n\n**Date\, Time and Location**\n\nAug 06\, 2026\n9:00 AM - 11:00 AM PST\n**[Online. Register for the Zoom!](https://voxel51.com/events/audio-and-ai-meetup-august-6-2026)**\n\n**Do Speech Models Actually Understand Speech? Evaluating Speech LLMs Under Realistic Spoken Instruction Conditions**\n\nSpeech Large Language Models (SLLMs) are increasingly capable\; but are we evaluating them the right way? Most benchmarks rely on text prompts\, yet real users interact with these systems through speech\, a modality that introduces noise\, disfluencies\, and stylistic variation that text simply doesn't capture.\nIn this talk\, we present findings from a systematic study across 11 tasks\, 12 languages\, and five prompt styles\, examining how prompt modality\, language\, and task type shape SLLM performance.\n\n*About the Speaker*\n\n[Maike Züfle ](https://voxel51.com/events/www.linkedin.com/in/maike-z%C3%BCfle)is a PhD student at the Karlsruhe Institute of Technology (KIT)\, working in Prof. Jan Niehues's group on interactive speech systems for more natural human–machine communication. Her research focuses on instruction-following speech models with speech as both input and output\, with a recent emphasis on full-duplex systems. Beyond her research\, she co-organises the instruction-following and speech translation metrics shared tasks at IWSLT. She is a 2026 Apple Scholar in AI/ML.\n\n**AI based Audio Forensics**\n\nIn this presentation\, attendees will discover several modules developed by Gradiant for the detection and analysis of synthetically generated or manipulated audio. The session will be delivered by one of the developers involved in the design and implementation of these technologies\, providing first-hand insight into their capabilities and underlying methodology.\n\nThe presentation will cover the traceability module\, which helps identify the origin of AI-generated content. It will also cover the segment detection tool\, designed to locate manipulated regions within an audio recording\, as well as the complete audio detection tool\, which assesses whether an entire recording has been synthetically generated.\n\n*About the Speaker*\n\n[Daniel Paniagua Ares ](https://voxel51.com/events/www.linkedin.com/in/daniel-paniagua-ares-96252918a)is a research engineer at Gradiant. Graduated in computer engineering from the FIC and with a master's degree in AI from the VIU.\n\n**Curating\, Searching\, and Evaluating Audio Datasets in FiftyOne**\n\nIn this talk\, we'll start with the ESC-50 environmental-sound dataset to show how FiftyOne represents audio: browsing clips in the tabular view\, rendering spectrograms directly in the sample grid with a custom renderer\, and turning sounds into searchable vectors with CLAP embeddings. Then we'll demo a similarity-search panel that lets you query an entire audio collection by example clip or a natural-language prompt to quickly find matching sounds.\n\nWe'll conclude with a live research problem: Audio Moment Retrieval from the DCASE 2026 Challenge\, where the goal is to localize the exact moment in a long recording that matches a text query. We'll frame this as temporal detection\, evaluate predictions\, and visualize ground-truth vs. predicted moments on an interactive timeline to intuitively expose model failure modes.\n\nAttendees will leave with a concrete blueprint and open code for applying visual data-centric AI practices to their own audio and multimodal datasets.\n\n*About the Speaker*\n\n[John Duncan](https://www.linkedin.com/in/john-a-duncan/) is a Machine Learning Engineer\, Customer Success at Voxel51. His research interests include vision\, LiDAR\, and audio perception for robots and intelligent systems.\n\nMer info: https://www.meetup.com/stockholm-ai-machine-learning-and-computer-vision-meetup/events/315389136/
LOCATION:Online event\, Stockholm
URL:https://www.meetup.com/stockholm-ai-machine-learning-and-computer-vision-meetup/events/315389136/
SOURCE:https://eveny.nu/event/aug-6-audio-and-ai-meetup-stockholm-20260806
END:VEVENT
END:VCALENDAR