652: A.I. Speech for the Speechless — with Jon Krohn (@JonKrohnLearns)

Topics covered
Popular Clips
Episode Highlights
AI Architecture
The innovative use of AI to aid speechless patients involves sophisticated deep learning architectures. explains how a machine vision model, developed by Dr. Arnib Haine at the University of Aachen, predicts speech from lip movements. This model utilizes convolutional neural networks (CNNs) and gated recurrent units (GRUs) to process video frames of patients' lips, enabling communication for those on mechanical ventilation 1.
The first stage of the model takes in video of lips moving and outputs a prediction of the audio that would be associated with that lip movement.
---
The two-stage model first estimates audio from lip movements and then converts this audio into text, significantly enhancing communication capabilities for ventilated patients 1.
Error Reduction
Efforts to reduce error rates in AI speech models are crucial for improving their accuracy and effectiveness. highlights the current 6.3% error rate of the model and the ongoing research to lower it by expanding the training dataset and exploring advanced deep learning architectures 2.
AI providing speech for the speechless having now prototyped their approach, the next step for these German clinical researchers is to expand their training dataset.
---
This research not only aims to enhance the model's precision but also to broaden its applicability to a wider range of patients, showcasing the practical and socially beneficial applications of AI 2.
Related Episodes

750: How AI is Transforming Science — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
SDS 464: A.I. vs Machine Learning vs Deep Learning — with Jon Krohn
Answers 383 questions
818: In Case You Missed It in August 2024 — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions

852: In Case You Missed It in December 2024 — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
SDS 438: Artificial General Intelligence — with Jon Krohn
Answers 383 questions
SDS 620: OpenAI Whisper: General-Purpose Speech Recognition — with @JonKrohnLearns
Answers 383 questions
832: The Anthropic CEO’s Techno-Utopia — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions

802: In Case You Missed It in June 2024 — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
740: Q*: OpenAI's Rumored AGI Breakthrough — with @JonKrohnLearns
Answers 383 questions
SDS 590: Artificial General Intelligence is Not Nigh (Part 2 of 2) — with Jon Krohn
Answers 383 questions
720: OpenAI’s DALL-E 3, Image Chat and Web Search — with Jon Krohn (@JonKrohnLearns)
Answers 383 questions
