Speech Recognition
cs.SD · 3 papers
cs.CL · 2024
Gemini 1.5: Unlocking Multimodal Long-Context
Synthetic Text Haystack: >99.7%Gemini Team, Petko Georgiev +1135
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning…
cs.CL · 2023
Gemini: A Family of Highly Capable Multimodal Models — Deep Dive
MMLU: 90.04%Gemini Team, Rohan Anil +1349
This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family…
eess.AS · 2022
Robust Speech Recognition via Large-Scale Weak Supervision — Whisper
Average relative error reduction: 55.2%Alec Radford, Jong Wook Kim +4
We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet. When scaled to 680,000 hours of…
Code· 102k★