Robert discusses the intriguing potential of large language models in enhancing causal analysis, revealing both their strengths and limitations when tested against existing benchmarks. While they excel in certain areas, questions remain about their effectiveness as standalone causal reasoners. Access to model weights and training data could be crucial for deeper insights into their capabilities and shortcomings.