Umberto Cappellazzo

Umberto Cappellazzo

Gen AI Research Engineer
NatWest Group AI Research
London, UK

About Me

I am a Gen AI Research Engineer at NatWest Group in London, working in the CAIRO group led by Prof. Maja Pantic, with Dr. Stavros Petridis as my manager. I'm bulding multimodal deepfake detectors and general-purpose audio learners at scale (millions of hours, from data to pre-training & post-training).

From 2025 to 2026 I was a Research Associate in the iBUG group at Imperial College London, where I still supervise students and keep ongoing projects running. I received my PhD in Information Engineering and Computer Science from the University of Trento in January 2025, defended cum laude, advised by Dr. Daniele Falavigna and Dr. Alessio Brutti.

My research is on speech, audio, and multimodal large language models: audio-visual speech recognition with LLMs, self-supervised audio representation learning, parameter-efficient fine-tuning, and understanding what multimodal models actually rely on.


Publications

For a full list, see my Google Scholar.

SSLAudio

Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners

U. Cappellazzo, X. Liu, S. Petridis, M. Pantic

Under review

AVSRInterpretabilityLLM

Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in AVSR

U. Cappellazzo, S. Petridis, M. Pantic

Interspeech 2026 (Long paper track)

AVSRLLM

VIB-AVSR: Variational Information Bottleneck for Noise-Robust LLM-Based AVSR

P. Arora, N. Singh, U. Cappellazzo, S. Petridis, M. Pantic

Interspeech 2026

PEFTAudioEfficiency

MambAdapter: Lightweight Mamba-Based Adapters for PEFT in Speech and Audio

S. Ali, U. Cappellazzo, M. Ravanelli

Interspeech 2026

AVSRLLMEfficiency

Omni-AVSR: Towards Unified Multimodal Speech Recognition with LLMs

U. Cappellazzo, X. Liu, P. Ma, S. Petridis, M. Pantic

2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

AVSRInterpretabilityLLM

Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs

Anand, U. Cappellazzo, S. Petridis, M. Pantic

2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

AVSRMoELLM

MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition

U. Cappellazzo, M. Kim, P. Ma, H. Chen, X. Liu, S. Petridis, M. Pantic

Neural Information Processing Systems 2025 (NeurIPS)

AVSRLLMEfficiency

Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs

U. Cappellazzo, M. Kim, S. Petridis

2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)

AVSRMoELLM

Scaling and Enhancing LLM-Based AVSR: A Sparse Mixture of Projectors Approach

U. Cappellazzo, M. Kim, S. Petridis, D. Falavigna, A. Brutti

Interspeech 2025

AVSRLLM

Large Language Models Are Strong Audio-Visual Speech Recognition Learners

U. Cappellazzo, M. Kim, H. Chen, P. Ma, S. Petridis, D. Falavigna, A. Brutti, M. Pantic

2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

PEFTAudio

Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers

U. Cappellazzo, D. Falavigna, A. Brutti, M. Ravanelli

2024 IEEE International Workshop on Machine Learning for Signal Processing (MLSP)

PEFTMoEAudio

Efficient Fine-Tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters

U. Cappellazzo, D. Falavigna, A. Brutti

Interspeech 2024

Continual LearningSLU

Continual Contrastive Spoken Language Understanding

U. Cappellazzo, E. Fini, M. Yang, D. Falavigna, A. Brutti, B. Raj

2024 Annual Meeting of the Association for Computational Linguistics (ACL Findings)

Continual LearningSLU

Evaluating and Improving Continual Learning in Spoken Language Understanding

M. Yang, X. Li, U. Cappellazzo, S. Watanabe, B. Raj

Interspeech 2024

Continual LearningAudio

Improving Continual Learning of Acoustic Scene Classification via Mutual Information Optimization

M. Yang, U. Cappellazzo, X. Li, S. Watanabe, B. Raj

2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

EfficiencyASR

Training Dynamic Models Using Early Exits for Automatic Speech Recognition on Resource-Constrained Devices

G. A. Wright, U. Cappellazzo, S. Zaiem, D. Raj, L. Ondel Yang, D. Falavigna, M. Ali, A. Brutti

2024 IEEE ICASSP Workshop

Continual LearningSLU

Sequence-Level Knowledge Distillation for Class-Incremental End-to-End Spoken Language Understanding

U. Cappellazzo, M. Yang, D. Falavigna, A. Brutti

Interspeech 2023

Continual LearningSLU

An Investigation of the Combination of Rehearsal and Knowledge Distillation in Continual Learning for Spoken Language Understanding

U. Cappellazzo, D. Falavigna, A. Brutti

Interspeech 2023


News

2026 Jun

Three papers accepted to Interspeech 2026 (one long, two regular): Dr. SHAP-AV, VIB-AVSR, MambAdapter. See you in Sydney.

2026 Mar

New paper: Dr. SHAP-AV, a study of modality contributions in AVSR at scale.

2026 Jan

Two papers accepted to ICASSP 2026: Omni-AVSR, and a study on attention sinks and massive activations in audio-visual LLMs.

2025 Sep

MoME accepted to NeurIPS 2025, unifying Matryoshka representation learning with sparse mixture-of-experts.

2025 Aug

Llama-MTSK accepted to ASRU 2025. See you in Honolulu.

2025 May

Llama-SMoP accepted to Interspeech 2025, a sparse mixture of projectors for LLM-based AVSR.

2025 Mar

Joined Imperial College London (iBUG) as a Research Associate.

2025 Jan

Defended my PhD cum laude at the University of Trento. [Dissertation] [Slides]


Experience

2026 Jul —

Gen AI Research Engineer, NatWest Group (CAIRO), LondonManager: Dr. Stavros Petridis. Group led by Prof. Maja Pantic.

2025 — 2026

Research Associate, Imperial College London (iBUG)Multimodal LLMs and self-supervised audio representation learning.

2024

Visiting Researcher, Imperial College London (iBUG)Nine-month visit on LLMs for AVSR — the work behind Llama-AVSR.

2023

JSALT Summer Workshop, Le MansFinite-state methods with modern neural architectures; early-exit techniques for CTC/MMI.


Education

2021 — 2025

PhD, Information Engineering and Computer Science, University of Trento"Efficient Knowledge Transfer and Adaptation for Speech and Beyond." Defended cum laude. Advised by Dr. Daniele Falavigna and Dr. Alessio Brutti.

2016 — 2019

MSc, Telecommunication Engineering, University of PadovaThesis: a deep-learning-based ECG delineator. Advised by Prof. Michele Rossi and Dr. Matteo Gadaleta.

2013 — 2016

BSc, Information Engineering, University of PadovaThesis: message authentication over an ideal or noisy channel. Advised by Prof. Nicola Laurenti.