Back

Research Area

Efficient, Mobile & Scalable AI Systems

We study efficient and scalable AI methods for mobile and product-scale settings.

Topics

Training Efficiency

Token dropping, masked pre-training, and two-stage schedules for diffusion transformers.

Sparse Architectures

Sparse-dense residual fusion, elastic latent interfaces, and adaptive compute allocation.

Token Pruning

Information-flow guided pruning for vision-language models under high compression ratios.

Few-step Generation

Flow matching, MeanFlow objectives, and curriculum strategies for one- and two-step inference.

Mobile AI

Quantization, distillation, and architecture search for on-device deep learning applications.

Systems & Infrastructure

Distributed training pipelines, large-scale infrastructure, and production ML deployment.

Publications

View all →

Explore another area