CFP: Evaluation Workshop on Speech and Language Technologies
Recent progress in speech and language technologies, increasingly driven by scale, with larger models, larger datasets, and more compute, has lead to impressive performance gains. However, these systems still have massive performance gaps that are glaringly obvious to human users, yet completely invisible to standard evaluation methods. As a result, evaluation has become a checklist exercise that reports incremental improvements on existing benchmarks rather than providing a meaningful assessment of what systems actually understand.