AI models
Choosing your LLM model
Compare the 15 available AI models — capability, context window, credits per reply, and ideal use cases.
The LLM model is the “brain” of your agent: it decides how questions are understood, how the knowledge base is used, and what tone the replies take. IperChat gives you access to 15 models from OpenAI and Anthropic — some are very fast and cheap, others reason deeply on complex problems. This page helps you pick the right one for your use case.
You set the model in the ops panel under “Agent configurations” (sidebar → Agents group). Open your origin and go to “Behavior” → “Model”. You can switch any time: the new model takes effect within 60 seconds on new conversations.
Comparison table
| Model | Provider | Context | Credits / msg | Best for |
|---|---|---|---|---|
| GPT-5.5 | OpenAI | 128k | 7 | Deepest reasoning, specialist scenarios |
| Claude Opus 5 | Anthropic | 200k | 5 | Most capable overall, long documents, complex multi-step conversations |
| Claude Opus 4.8 | Anthropic | 200k | 5 | Previous Opus generation, same price as Opus 5 |
| Claude Opus 4.7 | Anthropic | 200k | 5 | Long documents, extended thinking |
| Claude Opus 4.6 | Anthropic | 200k | 5 | Opus alternative, slightly older |
| GPT-5.6 Sol | OpenAI | 128k | 5 | OpenAI flagship, deep reasoning at lower cost than GPT-5.5 |
| GPT-5.4 | OpenAI | 128k | 3 | Complex conversations, high accuracy |
| Claude Sonnet 4.6 | Anthropic | 200k | 3 | Large KBs, lead qualification |
| GPT-5.6 Terra | OpenAI | 128k | 2 | Best quality / cost balance, high volume with reasoning |
| Claude Sonnet 5 | Anthropic | 200k | 2 | Near-Opus quality at Sonnet price, large knowledge bases |
| GPT-5.2 | OpenAI | 128k | 2 | Standard conversations with reasoning |
| GPT-5.6 Luna | OpenAI | 128k | 1 | Cheapest and fastest OpenAI reasoning model, maximum volume |
| GPT-4.1 Mini | OpenAI | 128k | 1 | Cheap generalist |
| GPT-4o | OpenAI | 128k | 1 | Battle-tested, multilingual |
| GPT-4o Mini | OpenAI | 128k | 1 | Maximum volume, minimum cost |
Tier 1 — Most capable (5–7 credits)
The premium models. Use them when accuracy matters more than per-message cost: regulated industries, specialist consulting, scenarios where a wrong answer is expensive.
GPT-5.5
OpenAI’s most capable model in our catalog. Configurable reasoning effort for problems that require multi-step thinking.
Best for:
- Legal, healthcare, financial consulting
- Multi-step reasoning over complex procedures
- Analytical summaries of technical documents
- Industries where a hallucination is costly
Claude Opus 5
Anthropic’s most capable model and the strongest in our catalog overall. Adaptive thinking with configurable reasoning effort: it reasons at length only when the question calls for it. No temperature setting. 200k-token context window.
Best for:
- The hardest cases: ambiguous questions, conversations with many constraints
- Very long documents and knowledge bases spread across dozens of files
- Complex multi-step conversations that need to keep track of everything said
- Regulated industries where reply quality is non-negotiable
Claude Opus 4.8
The previous Opus generation, at the same 5 credits as Opus 5. Adaptive thinking with configurable reasoning effort, no temperature setting, 200k context. Choose it if you tuned your prompts on 4.8 and want stable behaviour while you evaluate Opus 5.
Best for:
- Configurations already tuned on 4.8
- A/B comparisons against Opus 5 at equal cost
Claude Opus 4.7
Anthropic’s flagship with adaptive extended thinking: simple questions get an instant reply, complex ones trigger a “thinking” pass before answering. 200k-token context window — ideal for very large knowledge bases.
Best for:
- Questions that span many documents at once
- Analysis of contracts, reports, technical filings
- Conversations that need long-range context memory
- Natural, considered tone
Claude Opus 4.6
Previous Opus generation, kept available for setups already tuned on this version. Same characteristics as Opus 4.7 (200k context, extended thinking, 5 credits) with slightly less polished output.
Best for:
- Established configurations on 4.6 (for stability)
- A/B comparisons across Opus generations
GPT-5.6 Sol
The flagship of OpenAI’s GPT-5.6 family (released July 2026, knowledge cutoff February 2026). A reasoning model with configurable reasoning effort and no temperature setting. Deep reasoning at 5 credits — cheaper than GPT-5.5.
Best for:
- Deep reasoning on OpenAI when GPT-5.5’s cost is hard to justify
- Complex procedures and technical content that benefit from recent knowledge
- Specialist agents that need analytical depth at a lower price than GPT-5.5
Tier 2 — Balanced (2–3 credits)
The sweet spot for most production agents. Solid response quality, sustainable cost even at medium-to-high volume.
GPT-5.4
Recent OpenAI model with reasoning effort. Good quality/price tradeoff — great for conversations that need precision but not maximum analytical depth.
Best for:
- Complex conversations at moderate cost
- Questions requiring precise references to the KB
- Technical agents (IT, engineering, software)
Claude Sonnet 4.6
Sonnet is Anthropic’s workhorse: 200k context, extended thinking on demand, moderate cost. Default choice for agents that work over large knowledge bases.
Best for:
- Large knowledge bases (100+ documents)
- Lead qualification with detailed criteria
- Multilingual agents with a polished tone
- Booking and conversational flow management
GPT-5.6 Terra
The mid-size model of the GPT-5.6 family: a reasoning model with configurable reasoning effort, no temperature setting, knowledge cutoff February 2026. At 2 credits it is the best quality/cost balance in the catalog — our recommended default for most production agents.
Best for:
- Default choice for new agents
- High volume that still needs reasoning
- Lead qualification, booking, support with structured procedures
Claude Sonnet 5
Anthropic’s new Sonnet: close to Opus quality at Sonnet price. Adaptive thinking with configurable reasoning effort, no temperature setting, 200k context. At 2 credits it costs less than Sonnet 4.6 and replies better — the natural upgrade for large knowledge bases.
Best for:
- Large knowledge bases (100+ documents) on a budget
- Multilingual agents with a polished tone
- Anything you would run on Sonnet 4.6, at lower cost
GPT-5.2
Reasoning effort enabled at 2 credits per message. Cheaper than 5.4 while keeping reasoning capability.
Best for:
- Standard conversations with light reasoning
- Medium volume on a tight budget
Tier 3 — Fast & efficient (1 credit)
For high-volume use cases: standard customer support, repetitive FAQs, traffic deflection. Quick replies at low cost.
GPT-5.6 Luna
The smallest GPT-5.6 variant: the cheapest and fastest OpenAI reasoning model. Configurable reasoning effort, no temperature setting, knowledge cutoff February 2026. The best 1-credit option if you want reasoning at maximum volume.
Best for:
- Customer support over common FAQs
- Deflection of repetitive queries at high traffic
- Maximum volume on a controlled budget, without giving up reasoning
GPT-4.1 Mini
Classic OpenAI generalist. No reasoning, but supports temperature and top_p parameters if you want fine-grained control over response creativity.
Best for:
- Setups that need custom temperature
- Agents with very specific tone (creative or strict)
GPT-4o
OpenAI’s battle-tested multimodal model. Excellent multilingual support, natural writing quality.
Best for:
- Multilingual agents with international users
- Cases where you want a “stable” model that’s been around for months
- Existing setups that already work well
GPT-4o Mini
The mini version of 4o: the cheapest model in the OpenAI pool. Maximum volume at minimum cost.
Best for:
- Demos, prototypes, test environments
- Very high traffic where budget beats perfection
- First-line agents doing triage only
How to choose
A quick guide to narrow the choice based on your use case:
- Customer support, fast FAQs, high traffic → Tier 3: GPT-5.6 Luna.
- Lead qualification, booking, medium-complexity conversations → Tier 2: GPT-5.6 Terra or Claude Sonnet 5 (the default balanced pick).
- Regulated industries or specialist consulting (legal, healthcare, financial) → Tier 1: Claude Opus 5 for the hardest cases, or GPT-5.6 Sol / GPT-5.5 on OpenAI.
- Very large knowledge base (100+ documents, long contracts) → Claude family (200k context): Sonnet 5 or Opus 5.
- English agent on technical content → GPT-5.6 Terra or GPT-5.4.
- Demo, MVP, test traffic → GPT-4o Mini or GPT-5.6 Luna.