← Concept Library · Science & Technology
Science & Technology GS 3 In the news 4 times

Large Language Models (LLMs)

Technical Foundations

Large Language Models are deep neural networks trained on massive text corpora to understand and generate human language. They form the foundation of modern AI applications including chatbots, translation, summarisation, and code generation.

Key details
  • LLMs are typically based on the Transformer architecture (introduced in 2017), which uses self-attention mechanisms to process sequential data.
  • Model size is measured in parameters (learnable weights); larger models generally have greater capability but require more compute. Sarvam's 30B and 105B models represent medium and large-scale LLMs respectively.
  • Training involves pre-training on trillions of tokens (words/subwords) from text data, followed by fine-tuning for specific tasks.
  • Context length (how much text the model can process at once) is a key differentiator: 32K tokens for Sarvam-30B and 128K tokens for Sarvam-105B.
  • Open-source LLMs (such as Meta's Llama, Mistral, and now Sarvam) allow developers to inspect, modify, and deploy models without vendor lock-in, unlike proprietary models from companies like OpenAI or Google.
In the news

Tracked since February 18, 2026 · last seen February 25, 2026 · updates as the daily brief publishes

Related concepts
See it in today’s brief. Daily current affairs with every static concept explained in place.
Read the daily brief