AI models & usage
AI models & usage
Second Silicon runs each thread on a model you choose, or picks one for you from the providers you’ve configured. You can pin a model from the composer, or bring your own API key to use your existing provider quota.
Supported providers
The model picker beside the chat composer switches the model for the current thread. Settings → Models controls which models appear in that picker — disable the ones you don’t want to use.
Automatic selection
When you don’t pin a model, Second Silicon picks one from the providers you have a key for, in a cost-aware preference order. Pin a model from the composer whenever you want a specific one for the thread.
A separate per-thread reasoning effort setting controls how hard the model works before answering. Higher effort is more thorough but slower and more expensive.
Bring your own key (BYOK)
Add a provider API key to use your existing credits and keep full control over spend.
If no BYOK key is set for a provider, availability depends on the shared keys enabled for your workspace, and usage limits may apply during high-demand periods.
Token usage and cost
Every message in a debug thread shows three numbers: input tokens, output tokens, and an estimated USD cost. Input tokens are the context sent to the model. Output tokens are the model’s response.
Cost is an estimate based on the selected provider, model family, and current pricing assumptions in Second Silicon. It can differ from the provider’s bill, especially when a provider changes prices, applies discounts, or routes through a different underlying model.
Cost basis
Usage views label the pricing basis used for each estimate, such as Claude Sonnet (est.), GPT-5.4 mini (est.), Gemini Flash (est.), Kimi K2.6 via OpenRouter (est.), or GLM 5.2 via OpenRouter (est.). If a specific model has no exact match, Second Silicon falls back to the closest provider-level estimate.
Verify current pricing on the provider’s site before making cost decisions.