Automatic segmentation requires just the mask of an object from one frame, allowing the network to learn the object's appearance versus the background. However, challenges arise when similar objects appear in the scene, as illustrated by the example of two identical camels, which can confuse the network.