Back

Research Area

Language, Multimodal & Foundation Models

We conduct research on language, multimodal, and foundation models across text, images, and video.

Topics

Foundation Models & Scaling

Pretraining, scaling laws, and emergent capabilities of large language and multimodal models.

Text Embeddings

Decoder-based LLM embeddings, hierarchical token prepending, and long-document retrieval.

Multimodal Reasoning

In-context reasoning in diffusion models and vision-language alignment for generation.

Attention Mechanisms

Sparse attention, differential attention, and sink-free attention for long-context language modeling.

Multimodal Understanding

Vision-language grounding, cross-modal representation learning, and perception.

NLP & Information Extraction

Named entity recognition, structured moderation, and language understanding at scale.

Publications

View all →

Explore another area