Reconstructing the Wrong Winner: Choosing VAEs for Sign-Language Generation
TL;DR for operators A product team must choose one motion representation before spending substantially more compute training the generator that will use it. Reconstruction loss is a sensible first check: the representation must preserve the hand, face, and body information the product needs. The mistake is treating the cleanest reconstruction as proof that the downstream generator will learn best from it.1 ...