Models and Thinking Level
Which model the agent runs on, how hard it thinks, how credentials are resolved, and why a model choice can explain a behaviour you attributed to your code.
Something is not working. The agent is not using your tool. It is ignoring an instruction. It stops halfway. Before you go looking in your extension, check which model is running and what thinking level it is set to.
This is not a joke about blaming the model. It is a description of how the system is layered: model choice and thinking level sit underneath every mechanism in this book, and they change behaviour in ways that look exactly like bugs in the layer above.
Model, provider, credential
Three words that the documentation keeps distinct and that people routinely collapse.
A model generates Pi’s responses. A provider is the service or account Pi uses to reach that model. A credential is what makes the connection possible — a subscription, an API key, or something ambient.
The consequence is a specific and very common failure. Your custom model is
configured, the entry parses, the file reloads, and nothing appears in /model.
The reason is usually authentication: the picker shows models whose providers
have usable authentication. A model can load from models.json and still be
unavailable until Pi can resolve credentials for it.
Authenticating
Run /login, choose a provider, follow the prompts. Credentials go in
auth.json in the agent directory. /logout removes them for one provider.
Or supply an API key through the provider’s environment variable. This is the right choice in CI and anywhere Pi should not write credentials to disk.
When several sources are configured, Pi resolves them in this order:
flowchart TD
A["1 a runtime --api-key"] --> B["2 a stored auth.json credential"]
B --> C["3 an apiKey from models.json"]
C --> D["4 provider environment variables,<br/>or ambient cloud credentials"]
Provider extensions may define their own behaviour.
The first source that yields a credential wins, and the rest are never consulted.
That single rule explains a family of confusing symptoms. A stale auth.json will
beat the environment variable you just exported. An apiKey in models.json will
beat both. And if you changed an environment variable and the agent still cannot
authenticate, check that the variable is actually present in the process that
starts Pi — this is the usual cause of “it works in my shell but not in Pi”.
Wrong: “My key is in the environment, so Pi must be using it.”
Correct: “My key is in the environment, so Pi is using it unless a higher source in the chain resolved first.”
One warning belongs here. auth.json and any credential commands are private,
and project settings and extensions execute inside the Pi process once you trust
a project. Chapter 8 covers trust; chapter 36 covers what it costs.
Selecting a model
/model
opens a picker of available models with usable authentication. Ctrl+S on a
model saves it as the default for new sessions.
Ctrl+P cycles through available models. /scoped-models controls that cycle
and saves the selection, and you can configure model patterns through settings.
Two behaviours worth internalising:
A session records its model changes. Resuming the session restores the model and thinking level it was using, without changing your defaults for new sessions. So “it behaves differently now” may mean the session switched models two hours ago and never switched back.
The catalogue is not static. Pi starts with a bundled catalogue and can
overlay newer data from pi.dev; cached data stays available offline. Force a
refresh with pi update --models.
Thinking level
/thinking
selects the thinking level for the current model, and Pi limits the choices to
the levels that model supports. Ctrl+S saves the startup level.
The full set of settings values is "off", "minimal", "low", "medium",
"high", "xhigh", "max". The startup default is "medium".
Token budgets for minimal, low, medium and high come from
thinkingBudgets. modelThinkingLevels sets per-model startup levels keyed by
exact provider/modelId.
{
"defaultThinkingLevel": "medium",
"modelThinkingLevels": {
"anthropic/claude-sonnet-5": "high"
},
"thinkingBudgets": {
"high": 32000
}
}
Thinking blocks appear in the transcript. hideThinkingBlock hides them. This
is a display setting, and it matters for extension authors: what you see in the
terminal is not necessarily what was stored.
This is where agent behaviour gets harder to debug, not easier. A model at
off and a model at high are two different systems with respect to how they
approach a task. If your extension’s tool is being called in a way you did not
expect, confirm the thinking level before you change the tool.
Tools, startup, and defaults
Three settings govern what starts where:
| Setting | What it does |
|---|---|
defaultProvider |
Startup provider; automatic by default |
defaultModel |
Startup model ID; automatic by default |
enabledModels |
Model patterns used for startup selection and cycling |
enabledModels defaults to all available models and uses the same format as
the --models flag.
Prompt caching, and why it appears here
Caching is a model-level concern, and it shows up in agent behaviour more often than you would expect.
promptCache declares the provider’s best-effort cache lifetime, in seconds,
for the short or long retention tier:
{ "id": "claude-sonnet-5", "promptCache": { "short": 300, "long": 3600 } }
Choose the conservative end of any published range. A model without a lifetime for the active tier is not eligible for cache warming.
cacheWarming is "off", "streaming" or "idle", defaulting to
"streaming". Warming runs only when the model declares a cache lifetime and
Pi estimates at least $0.05 in avoided cache-miss cost. Refresh usage counts
toward session totals but does not enter model context — which is a detail to
remember when your context accounting and your billing disagree.
showCacheMissNotices surfaces significant cache misses, successful warming,
compaction usage and provider recovery. If you are seeing unexplained latency
or unexplained token counts, turn it on before you instrument anything.
Models that are not chat models
Two categories exist that do not appear in /model, and knowing they exist
prevents a class of confusion.
Classifier models do not chat. They answer typed questions about JSON state: choose one of several options, answer yes or no, or produce a score, each with probabilities.
Image models generate images from a prompt and optional input images.
Neither appears in the /model picker. Both are reached through the codemode
tool, which a built-in extension registers inactive. Enable it yourself with:
{ "defaultTools": ["+codemode"] }
Extensions reach them directly instead, without codemode, through
ctx.modelRegistry.classify() and ctx.modelRegistry.generateImages().
If you have been looking for an image model in /model and it is not there,
this is why, and it is not a bug.
Endpoints you can talk to
When Pi does not already support the provider or endpoint you need, there are three rungs, in increasing order of effort:
-
A compatible endpoint. If the endpoint speaks an API Pi already supports — most Ollama, LM Studio, vLLM, SGLang and proxy deployments — add it to
models.jsonin the agent directory.{ "providers": { "ollama": { "baseUrl": "http://localhost:11434/v1", "api": "openai-completions", "apiKey": "ollama", "models": [ { "id": "qwen2.5-coder:7b" } ] } } }The dummy key makes the model available; Ollama ignores it. For an authenticated endpoint,
apiKeyand header values accept$NAMEor${NAME}environment interpolation, a literal, or a leading!command. Commands run at request time and are not cached by Pi.Opening
/modelreloads the file. -
A custom provider extension, when the provider needs custom streaming, model discovery or authentication behaviour.
-
A virtual model, which registers under your own name and routes each request to whatever you decide — including a classifier. Chapter 36 returns to this, because “route it somewhere else” and “it may do that” are different decisions.
One warning from the documentation that is worth repeating because it is a common and expensive mistake: describe only verified differences in an endpoint’s request or response behaviour in your compatibility settings. Do not enable them because the endpoint advertises OpenAI or Anthropic compatibility.
What this chapter is for
The practical content is the tables and the precedence order. The idea is smaller and more useful:
model choice affects what is possible
thinking level affects how it approaches the task
credentials affect whether anything happens at all
cache behaviour affects latency and reported usage
None of these are your extension. All of them can produce a symptom you will otherwise debug for an hour in the wrong layer. When something is behaving impossibly, check this chapter’s list before you check your code.
Next
You now know what the loop runs on. What it runs with is the next constraint, and it is finite: every request competes for a fixed window, and your instructions are one line item among several rather than the whole container. Knowing that turns “the agent got worse” into arithmetic you can read.
Chapter 4: the context window as a budget.