One оf yоur SEAS 8525 cоlleаgues stаtes: "ViT must use а completely different architecture than the NLP transformer because images and text are totally different modalities." Which response most accurately corrects his statement?