Published Aug 5, 2016

Machine Learning on Images with Noisy Human-centric Labels

Ishan Misra delves into the transformative use of large, uncurated datasets to refine image classification by addressing noisy data and human reporting bias, offering innovative strategies for enhancing model accuracy and interpretability in non-curated datasets.
Episode Highlights
Data Skeptic logo

Popular Clips

Episode Highlights

  • Optimization

    The optimization techniques used in image classification are crucial for improving model performance. explains that they initially trained a model to predict human-like labels, holding relevance constant, before training for relevance and visual presence simultaneously. This approach, which involves alternating optimization, led to significant improvements. On non-curated datasets, they doubled the performance of baseline systems, achieving over 5% gains in average precision across 1000 tags 1.

    We have gains of about more than 5% in average precision over 1000 tags.

    ---

    For curated datasets, the gains were about 3% in average precision, demonstrating the effectiveness of their method.

       

    Presence vs. Relevance

    Separating visual presence from relevance in image classification offers distinct advantages. highlights that not separating these factors leads to inconsistent labels, making it difficult for models to mimic human perception accurately. By solving for these factors separately, models can achieve visually correct predictions rather than merely mimicking human labels 2.

    You will never be able to learn a model that gives you visually correct predictions if you never separate the two factors out.

    ---

    This separation also allows for better interpretation of visual concepts, as demonstrated by a 12x12 grid output that localizes visual concepts within images 3.

       

    Relevance

    Optimizing relevance scores in image classification presents unique challenges. discusses how a vanilla classification network struggled with loss reduction due to inconsistent labeling, such as predicting 'yellow' for images with minimal yellow presence. By providing the model with extra capacity, they enabled it to assign high visual presence but low relevance to certain features, effectively reducing training loss 4.

    The network is expected to predict yellow, but for this image which has yellow all over, you're expecting the network not to predict yellow.

    ---

    This approach addresses the inconsistency in learning, allowing for more accurate predictions.

Related Episodes