This article considers Sonnet 5.5 and Opus 5.5 for coding workflows. The recommendations are vendor guidance; the evaluation method below is my proposal, with no independent measurements presented here.
Choosing a model
Anthropic describes model selection as a balance between capabilities, speed, cost and reasoning effort. Its general guide offers two strategies:
- Efficiency first: start with Haiku 4.5, test the use case and upgrade when results reveal capability gaps.
- Capability first: start with Opus 5.5 for complex work, evaluate, then consider reducing effort or using a less expensive model.
The same guide says most workloads start with Opus 5.5 and identifies Fable 5.1 for the highest available capability. Fable is outside this narrower Sonnet–Opus comparison. (Anthropic, n.d., Choosing the right model (opens in a new tab)).
Osmani’s narrower Sonnet 5.5 guide recommends:
| Workload | Recommended starting model |
|---|---|
| Well-scoped everyday coding, with clear requirements and checkable results | Sonnet 5.5 |
| Complex work requiring careful judgement | Opus 5.5 |
These are starting recommendations with different emphases, rather than one universal selection rule. (Osmani, 2026, Building with Claude Sonnet 5.5 (opens in a new tab)).
Tuning Sonnet’s effort
Choosing a model and choosing its reasoning-effort setting are separate decisions. For Sonnet 5.5 agentic coding and multistep tool use, Osmani recommends:
| Task within this workload | Starting effort |
|---|---|
| Well-specified task | Medium |
| Harder or longer task | High |
Outside the guide’s agentic and latency-sensitive exceptions, it recommends starting at high. It recommends xhigh or max only where evaluations demonstrate a quality gain. Medium is therefore not its recommendation for every task. (Osmani, 2026, Building with Claude Sonnet 5.5 (opens in a new tab)).
Record the application or API surface and set effort explicitly where possible. Osmani reports different Sonnet defaults: high on Claude Platform and medium in Claude Code. Comparing unspecified defaults can therefore compare different configurations unintentionally. (Osmani, 2026).
Evaluating a configuration
Anthropic recommends use-case-specific benchmark tests with actual prompts and data. Compare accuracy, response quality and handling of edge cases, then weigh performance against cost. (Anthropic, n.d., Choosing the right model (opens in a new tab)).
My proposed method is to choose a representative set of tasks, including difficult cases, and define acceptance criteria before comparing configurations. Set a time limit, usage-cost limit and maximum number of retries for each run. A run that does not meet the criteria within those limits is a failure; retain its cost and time in the record.
Run each task several times with each configuration, choosing the same number of runs in advance. Keep the starting prompts, tools, supplied context and acceptance checks consistent. Record any human intervention and use the same rules for allowing corrections or retries. Treat a model or effort change as a different configuration.
Start each run from a fresh session and the same repository snapshot. Record the test environment, dependency versions and cache conditions. Alternate configuration order to reduce timing effects, retain prompts, outputs and usage records, and assess results without model labels where practical.
| Record | What to capture |
|---|---|
| Configuration | Exact model version, effort setting, application or API surface, client version and other generation settings |
| Task and criteria | Task identifier, starting prompt, tools, context and acceptance checks |
| Limits | Maximum elapsed time, usage cost and retries per run |
| Result | Pass or failure, supporting checks and reason for failure or stopping |
| Usage cost | Recorded model usage cost for the whole run, including retries and failed attempts |
| Elapsed time | Wall-clock time from start to acceptance or stopping, including waits |
| Human review time | Active time spent checking and correcting, recorded separately from elapsed time |
| Intervention | Corrections, retries and model switches, with their additional cost and time |
Report success rate as accepted runs divided by all runs, alongside the number of tasks and repetitions. Show results by task type as well as overall, so a strong average does not conceal a recurring failure. Report elapsed time and active review time for both successful and failed runs; do not assign an unsuccessful run a time to an acceptable result.
For usage cost per accepted result, divide the total usage cost of all runs, including failures, by the number of accepted results. If none succeed, report that no accepted result was obtained rather than quoting a cost per success. Keep human review time visible alongside usage cost; if converting it to money, state the hourly rate used. These records support a decision about the tested workload, not a claim about every task.
Illustrative example — hypothetical figures, not model measurements. Select ten Python bug-fixing tasks covering ordinary inputs and edge cases. Require the existing tests and predefined acceptance checks to pass, with no unrelated changes. Run each task three times per configuration from a clean copy, allowing ten minutes, US$2 in model usage and one retry per run.
Suppose configuration A accepts 24 of 30 runs, consumes US$18 across all runs and requires 90 minutes of active review. Configuration B accepts 27, consumes US$24 and requires 60 minutes of review. Acceptance rates are 80% and 90%; usage costs per accepted result are US$0.75 and approximately US$0.89. At an assumed review rate of US$30 per hour, combined usage and review costs per accepted result become US$2.63 and US$2.00. Include failed runs in every total. Report completion times separately and check both configurations against the latency requirement. These small hypothetical samples establish no model ranking.
The cost of an acceptable result
Artificial Analysis reports an Intelligence Index score of 56 and US$7.60 per index task for Sonnet 5.5 at max effort. Its index also lists Opus 5.5 at xhigh with a score of 56. These are producer-reported measurements, not evaluations reproduced here. The Sonnet report notes that testing used a pre-release deployment with a structured-output bug. (Artificial Analysis, 2026a (opens in a new tab); 2026b (opens in a new tab)).
Cost per index task is a weighted benchmark average calculated using input, cache-hit, cache-write, reasoning and answer token prices. It is not the cost of successfully completing an arbitrary real-world task. Equal aggregate scores do not establish equal strengths or failure rates on individual tasks. (Artificial Analysis, 2026b).
If the workflow switches models, measure the additional latency, context processing and review effort.
A lower reasoning-effort setting does not, by itself, establish a lower cost per acceptable result. Choose the configuration that meets your quality and latency requirements at the lowest observed total cost. Include failed attempts, retries and review time, stating any rate used to value that time. Assess the consequences of undetected errors separately, with explicit assumptions about their likelihood and impact.