Skip to main content
IperChat
Try free demo IT
Log in

AI models

Choosing your LLM model

Compare the 15 available AI models — capability, context window, credits per reply, and ideal use cases.

The LLM model is the “brain” of your agent: it decides how questions are understood, how the knowledge base is used, and what tone the replies take. IperChat gives you access to 15 models from OpenAI and Anthropic — some are very fast and cheap, others reason deeply on complex problems. This page helps you pick the right one for your use case.

You set the model in the ops panel under “Agent configurations” (sidebar → Agents group). Open your origin and go to “Behavior” → “Model”. You can switch any time: the new model takes effect within 60 seconds on new conversations.

Comparison table

ModelProviderContextCredits / msgBest for
GPT-5.5OpenAI128k7Deepest reasoning, specialist scenarios
Claude Opus 5Anthropic200k5Most capable overall, long documents, complex multi-step conversations
Claude Opus 4.8Anthropic200k5Previous Opus generation, same price as Opus 5
Claude Opus 4.7Anthropic200k5Long documents, extended thinking
Claude Opus 4.6Anthropic200k5Opus alternative, slightly older
GPT-5.6 SolOpenAI128k5OpenAI flagship, deep reasoning at lower cost than GPT-5.5
GPT-5.4OpenAI128k3Complex conversations, high accuracy
Claude Sonnet 4.6Anthropic200k3Large KBs, lead qualification
GPT-5.6 TerraOpenAI128k2Best quality / cost balance, high volume with reasoning
Claude Sonnet 5Anthropic200k2Near-Opus quality at Sonnet price, large knowledge bases
GPT-5.2OpenAI128k2Standard conversations with reasoning
GPT-5.6 LunaOpenAI128k1Cheapest and fastest OpenAI reasoning model, maximum volume
GPT-4.1 MiniOpenAI128k1Cheap generalist
GPT-4oOpenAI128k1Battle-tested, multilingual
GPT-4o MiniOpenAI128k1Maximum volume, minimum cost

Tier 1 — Most capable (5–7 credits)

The premium models. Use them when accuracy matters more than per-message cost: regulated industries, specialist consulting, scenarios where a wrong answer is expensive.

GPT-5.5

OpenAI’s most capable model in our catalog. Configurable reasoning effort for problems that require multi-step thinking.

Best for:

  • Legal, healthcare, financial consulting
  • Multi-step reasoning over complex procedures
  • Analytical summaries of technical documents
  • Industries where a hallucination is costly

Claude Opus 5

Anthropic’s most capable model and the strongest in our catalog overall. Adaptive thinking with configurable reasoning effort: it reasons at length only when the question calls for it. No temperature setting. 200k-token context window.

Best for:

  • The hardest cases: ambiguous questions, conversations with many constraints
  • Very long documents and knowledge bases spread across dozens of files
  • Complex multi-step conversations that need to keep track of everything said
  • Regulated industries where reply quality is non-negotiable

Claude Opus 4.8

The previous Opus generation, at the same 5 credits as Opus 5. Adaptive thinking with configurable reasoning effort, no temperature setting, 200k context. Choose it if you tuned your prompts on 4.8 and want stable behaviour while you evaluate Opus 5.

Best for:

  • Configurations already tuned on 4.8
  • A/B comparisons against Opus 5 at equal cost

Claude Opus 4.7

Anthropic’s flagship with adaptive extended thinking: simple questions get an instant reply, complex ones trigger a “thinking” pass before answering. 200k-token context window — ideal for very large knowledge bases.

Best for:

  • Questions that span many documents at once
  • Analysis of contracts, reports, technical filings
  • Conversations that need long-range context memory
  • Natural, considered tone

Claude Opus 4.6

Previous Opus generation, kept available for setups already tuned on this version. Same characteristics as Opus 4.7 (200k context, extended thinking, 5 credits) with slightly less polished output.

Best for:

  • Established configurations on 4.6 (for stability)
  • A/B comparisons across Opus generations

GPT-5.6 Sol

The flagship of OpenAI’s GPT-5.6 family (released July 2026, knowledge cutoff February 2026). A reasoning model with configurable reasoning effort and no temperature setting. Deep reasoning at 5 credits — cheaper than GPT-5.5.

Best for:

  • Deep reasoning on OpenAI when GPT-5.5’s cost is hard to justify
  • Complex procedures and technical content that benefit from recent knowledge
  • Specialist agents that need analytical depth at a lower price than GPT-5.5

Tier 2 — Balanced (2–3 credits)

The sweet spot for most production agents. Solid response quality, sustainable cost even at medium-to-high volume.

GPT-5.4

Recent OpenAI model with reasoning effort. Good quality/price tradeoff — great for conversations that need precision but not maximum analytical depth.

Best for:

  • Complex conversations at moderate cost
  • Questions requiring precise references to the KB
  • Technical agents (IT, engineering, software)

Claude Sonnet 4.6

Sonnet is Anthropic’s workhorse: 200k context, extended thinking on demand, moderate cost. Default choice for agents that work over large knowledge bases.

Best for:

  • Large knowledge bases (100+ documents)
  • Lead qualification with detailed criteria
  • Multilingual agents with a polished tone
  • Booking and conversational flow management

GPT-5.6 Terra

The mid-size model of the GPT-5.6 family: a reasoning model with configurable reasoning effort, no temperature setting, knowledge cutoff February 2026. At 2 credits it is the best quality/cost balance in the catalog — our recommended default for most production agents.

Best for:

  • Default choice for new agents
  • High volume that still needs reasoning
  • Lead qualification, booking, support with structured procedures

Claude Sonnet 5

Anthropic’s new Sonnet: close to Opus quality at Sonnet price. Adaptive thinking with configurable reasoning effort, no temperature setting, 200k context. At 2 credits it costs less than Sonnet 4.6 and replies better — the natural upgrade for large knowledge bases.

Best for:

  • Large knowledge bases (100+ documents) on a budget
  • Multilingual agents with a polished tone
  • Anything you would run on Sonnet 4.6, at lower cost

GPT-5.2

Reasoning effort enabled at 2 credits per message. Cheaper than 5.4 while keeping reasoning capability.

Best for:

  • Standard conversations with light reasoning
  • Medium volume on a tight budget

Tier 3 — Fast & efficient (1 credit)

For high-volume use cases: standard customer support, repetitive FAQs, traffic deflection. Quick replies at low cost.

GPT-5.6 Luna

The smallest GPT-5.6 variant: the cheapest and fastest OpenAI reasoning model. Configurable reasoning effort, no temperature setting, knowledge cutoff February 2026. The best 1-credit option if you want reasoning at maximum volume.

Best for:

  • Customer support over common FAQs
  • Deflection of repetitive queries at high traffic
  • Maximum volume on a controlled budget, without giving up reasoning

GPT-4.1 Mini

Classic OpenAI generalist. No reasoning, but supports temperature and top_p parameters if you want fine-grained control over response creativity.

Best for:

  • Setups that need custom temperature
  • Agents with very specific tone (creative or strict)

GPT-4o

OpenAI’s battle-tested multimodal model. Excellent multilingual support, natural writing quality.

Best for:

  • Multilingual agents with international users
  • Cases where you want a “stable” model that’s been around for months
  • Existing setups that already work well

GPT-4o Mini

The mini version of 4o: the cheapest model in the OpenAI pool. Maximum volume at minimum cost.

Best for:

  • Demos, prototypes, test environments
  • Very high traffic where budget beats perfection
  • First-line agents doing triage only

How to choose

A quick guide to narrow the choice based on your use case:

  • Customer support, fast FAQs, high traffic → Tier 3: GPT-5.6 Luna.
  • Lead qualification, booking, medium-complexity conversations → Tier 2: GPT-5.6 Terra or Claude Sonnet 5 (the default balanced pick).
  • Regulated industries or specialist consulting (legal, healthcare, financial) → Tier 1: Claude Opus 5 for the hardest cases, or GPT-5.6 Sol / GPT-5.5 on OpenAI.
  • Very large knowledge base (100+ documents, long contracts) → Claude family (200k context): Sonnet 5 or Opus 5.
  • English agent on technical contentGPT-5.6 Terra or GPT-5.4.
  • Demo, MVP, test trafficGPT-4o Mini or GPT-5.6 Luna.

Next steps