Sequence NLL Loss

Matt and Sergey discuss the differences between sequence and token level NLL loss, highlighting how the gradient computation varies for each token in a sequence. They delve into the nuances of probability assignment and normalization, shedding light on the intricacies of modeling language sequences.