Machine Learning on Images with Noisy Human-centric Labels

Topics covered
Popular Clips
Episode Highlights
Optimization
The optimization techniques used in image classification are crucial for improving model performance. explains that they initially trained a model to predict human-like labels, holding relevance constant, before training for relevance and visual presence simultaneously. This approach, which involves alternating optimization, led to significant improvements. On non-curated datasets, they doubled the performance of baseline systems, achieving over 5% gains in average precision across 1000 tags 1.
We have gains of about more than 5% in average precision over 1000 tags.
---
For curated datasets, the gains were about 3% in average precision, demonstrating the effectiveness of their method.
  Â
Presence vs. Relevance
Separating visual presence from relevance in image classification offers distinct advantages. highlights that not separating these factors leads to inconsistent labels, making it difficult for models to mimic human perception accurately. By solving for these factors separately, models can achieve visually correct predictions rather than merely mimicking human labels 2.
You will never be able to learn a model that gives you visually correct predictions if you never separate the two factors out.
---
This separation also allows for better interpretation of visual concepts, as demonstrated by a 12x12 grid output that localizes visual concepts within images 3.
  Â
Relevance
Optimizing relevance scores in image classification presents unique challenges. discusses how a vanilla classification network struggled with loss reduction due to inconsistent labeling, such as predicting 'yellow' for images with minimal yellow presence. By providing the model with extra capacity, they enabled it to assign high visual presence but low relevance to certain features, effectively reducing training loss 4.
The network is expected to predict yellow, but for this image which has yellow all over, you're expecting the network not to predict yellow.
---
This approach addresses the inconsistency in learning, allowing for more accurate predictions.
Related Episodes


Fooling Computer Vision
Answers 383 questions

Computer Vision is Not Perfect
Answers 383 questions

Machine Learning Done Wrong
Answers 383 questions

Visual Illusions Deceiving Neural Networks
Answers 383 questions

Transfer Learning
Answers 383 questions

Black Boxes Are Not Required
Answers 383 questions

Visualization and Interpretability
Answers 383 questions

Drug Discovery with Machine Learning
Answers 383 questions

Easily Fooling Deep Neural Networks
Answers 383 questions

Understanding Neural Networks
Answers 383 questions
Human vs Machine Transcription
Answers 383 questions

Annotator Bias
Answers 383 questions
[MINI] Noise!!
Answers 383 questions

Applied Data Science in Industry
Answers 383 questions

Interpretable One Shot Learning
Answers 383 questions
