Alessio explains the limitations of RNNs in parallelization due to sequential dependencies, while Shawn highlights the efficiency of transformers in processing variable-length sequences in a single pass during training. Inference also incurs additional costs with transformers compared to RNNs.