arXiv Artificial Intelligence

Federated Multilingual Speech-LLMs: Architecture and Aggregation Strategy Benchmarking

Federated Multilingual Speech-LLMs: Architecture and Aggregation Strategy Benchmarking

Quick summary

arXiv:2609.23825v1 Announce Type: cross Abstract: We present a comprehensive benchmark of Federated Learning (FL) for multilingual Automatic Speech Recognition (ASR), evaluating four Speech-LLM architectures on the Multilingual LibriSpeech dataset. We compare FedAvg and FedProx across frozen and unfrozen encoder configurations, demonstrating that optimized learning rates are critical for performance. Specifically, independently tuning the learning rates for the speech encoder, connector, and decoder yields the lowest error rates, with full three-component adaptation (LoRA for encoder and decod

Key takeaways

  • arXiv:2609.23825v1 Announce Type: cross Abstract: We present a comprehensive benchmark of Federated Learning (FL) for multilingual Automatic Speech Recognition (ASR), evaluating four Speech-LLM architectures on the Multilingual LibriSpeech dataset.
  • We compare FedAvg and FedProx across frozen and unfrozen encoder configurations, demonstrating that optimized learning rates are critical for performance.
  • Specifically, independently tuning the learning rates for the speech encoder, connector, and decoder yields the lowest error rates, with full three-component adaptation (LoRA for encoder and decod

Why it matters

“Federated Multilingual Speech-LLMs: Architecture and Aggregation Strategy Benchmarking” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗