Claude model tiers की त्वरित, व्यावहारिक गाइड। Routing blueprints, prompting patterns और token-budget playbook वाले पूर्ण deep-dive के लिए पढ़ें बड़ा article.
claude-haiku-4-5 तेज़ और सस्ते intent काम के लिए, claude-sonnet-5 सर्वोत्तम speed/quality balance के लिए, claude-opus-4-8 जटिल agentic coding और enterprise काम के लिए, और आरक्षित रखें claude-fable-5 (और gated claude-mythos-5) वास्तव में कठिन, long-horizon jobs के लिए जो घंटों या दिनों लेते हैं।
एक नज़र में model lineup
| मॉडल | सबसे अच्छा | Context window | Max output | Pricing (input / output) | Thinking behavior |
|---|---|---|---|---|---|
Claude Fable 5 (claude-fable-5) |
Long-running agents और सबसे कठिन unresolved समस्याएँ | 1M tokens | 128K tokens | $10 / $50 per MTok | Adaptive thinking always on |
Claude Opus 4.8 (claude-opus-4-8) |
जटिल agentic coding और enterprise काम | 1M tokens | 128K tokens | $5 / $25 per MTok | Adaptive thinking always on |
Claude Sonnet 5 (claude-sonnet-5) |
सर्वोत्तम speed/quality balance | 1M tokens | 128K tokens | $3 / $15 per MTok | Adaptive thinking always on |
Claude Haiku 4.5 (claude-haiku-4-5-20251001) |
सस्ते, high-volume tasks के लिए सबसे तेज़ model | 200K tokens | 64K tokens | $1 / $5 per MTok | Extended thinking available (adaptive off) |
Claude Mythos 5 (claude-mythos-5) एक gated model है जो Fable 5 जैसी ही specs और pricing साझा करता है, लेकिन Fable 5 वाले safety classifiers शामिल नहीं करता। उपलब्धता Project Glasswing के माध्यम से अनुमोदित partners तक सीमित है।
सही model चुनें (त्वरित नियम)
- Haiku 4.5: routing, classification, extraction, और कोई भी workflow जहाँ तेज़ जवाब चाहिए और उच्च output-token budgets वहन नहीं किए जा सकते।
- Sonnet 5: दैनिक coding और document काम जहाँ Opus/Fable कीमत के बिना मजबूत गुणवत्ता चाहिए।
- Opus 4.8: जटिल agentic coding, enterprise analysis, और कठिन debugging जहाँ लंबे, अधिक autonomous काम की जरूरत हो।
- Fable 5: वे कठिन, long-horizon समस्याएँ जो घंटों या दिनों लेती हैं, खासकर जब model को लक्ष्य रखना, sub-tasks सौंपना और सुसंगत रहना हो।
Effort और thinking: लागत पर प्रभाव
API पर, effort Fable 5 और Mythos 5 पर intelligence, latency और cost के बीच trade-off का प्राथमिक नियंत्रण है। Opus 4.8 और Sonnet 5 के लिए आप effort स्पष्ट रूप से भी सेट कर सकते हैं जब अलग cost/latency profile चाहिए।
व्यावहारिक अर्थ: हर request को maximum effort पर डिफ़ॉल्ट करना अक्सर output tokens पर बजट खर्च करने का सबसे तेज़ तरीका है।
Token strategy जो सभी tiers पर काम करे
Claude output tokens महँगा हिस्सा हैं। एक सरल नियम है: output length सीमित करें, TLDR-first के लिए prompt, और उपयोग करें prompt caching ताकि stable prefixes (system instructions और tool schemas) पुन: उपयोग हों।
- सबसे सस्ता model उपयोग करें जो पूरा कर सके। एक router (Haiku -> Sonnet/Opus -> Fable) अक्सर गुणवत्ता ऊँची रखते हुए लागत नाटकीय रूप से काट देता है।
- Output limits सेट करें। Production में हमेशा max output / token cap सेट करें। असीमित उत्तर budgeting accident हैं।
- Summaries रणनीतिक रूप से उपयोग करें। हर turn पर पूरे transcripts भेजने के बजाय, एक compact “lessons” memory रखें और केवल उसे अपडेट करें।
एक सरल router blueprint
- Intent और अनुमानित complexity classify करें
claude-haiku-4-5. - “Normal hard” काम भेजें
claude-sonnet-5(याclaude-opus-4-8जब आप heavier agentic coding की उम्मीद करें)। - बढ़ाएँ
claude-fable-5केवल जब task स्पष्ट रूप से long-horizon हो या निचले tiers पर fail हो चुका हो। - यदि Fable 5 safety classifiers के माध्यम से मना करे, Opus 4.8 पर वापस जाएँ (API configured होने पर fallbacks के साथ refiring support करता है)।