Research Area
We study efficient and scalable AI methods for mobile and product-scale settings.
Topics
Token dropping, masked pre-training, and two-stage schedules for diffusion transformers.
Sparse-dense residual fusion, elastic latent interfaces, and adaptive compute allocation.
Information-flow guided pruning for vision-language models under high compression ratios.
Flow matching, MeanFlow objectives, and curriculum strategies for one- and two-step inference.
Quantization, distillation, and architecture search for on-device deep learning applications.
Distributed training pipelines, large-scale infrastructure, and production ML deployment.
Explore another area