Multimodal AI: computer vision or audio/speech processing, student's choice
This course deepens specialization in one of two multimodal tracks: Computer Vision (detection, segmentation, vision-language models) or Audio/Speech (ASR, TTS, speaker diarization, multimodal audio models). It covers both modality-specific architectural choices and general multimodal fusion approaches. The course prepares students for the research track (s3-ai-research) and applied projects (s3-applied-ai).