Andrew emphasizes the importance of local inference for devices like cars and phones, highlighting that the future of inference will increasingly involve machine-to-machine interactions rather than human-readable outputs. He warns against the pitfalls of latency stacking and discusses the evolving complexity of enterprise workflows that leverage collections of inference.