No Distillation Needed
Single-Pass Real-Time Talking Heads via Acausal Noise Shaping

Yu Han, Dejan Markovic, Alexander Richard, Wojciech Zielonka, Akshay Venkatesh, Cheng-hsin Wuu and Michael Zöllhöfer

Meta Reality Labs

FaceGAN teaser: streaming audio to rendered talking heads in one forward pass; acausal noise shaping yields smooth, expressive motion.

FaceGAN Method Comparison

About. Supplementary qualitative results for No Distillation Needed: Single-Pass Real-Time Talking Heads via Acausal Noise Shaping. FaceGAN is a streaming single-pass GAN that maps speech to facial expression and head pose in one forward pass per frame, using latency-free acausal noise shaping to produce smooth, rich motion without diffusion's iterative sampling cost. Each row compares all methods on one test clip driven by the same Ground Truth audio.