Research Area
We conduct research on language, multimodal, and foundation models across text, images, and video.
Topics
Pretraining, scaling laws, and emergent capabilities of large language and multimodal models.
Decoder-based LLM embeddings, hierarchical token prepending, and long-document retrieval.
In-context reasoning in diffusion models and vision-language alignment for generation.
Sparse attention, differential attention, and sink-free attention for long-context language modeling.
Vision-language grounding, cross-modal representation learning, and perception.
Named entity recognition, structured moderation, and language understanding at scale.
Explore another area