Model Evaluation: Fable vs. Opus for Agentic Engineering Work
Executive Summary
Over a multi-day engineering evaluation, Fable showed strong performance as a fast implementation model: it moved quickly, handled scoped tasks well, and responded effectively to precise review feedback. However, we did not find enough evidence to conclude that Fable is broadly superior to Opus for high-judgment work.
The best operating model is not “Fable or Opus.” It is a split workflow:
Use Fable for implementation, iteration, proof repair, and tightly scoped engineering tasks.
Use Opus for independent review, architecture, risk analysis, public-claim validation, and final acceptance gates.
Conclusion
Fable is well suited for fast builder work under clear constraints. Opus remains better suited for judgment-heavy review and high-risk decisions.
What We Evaluated
We compared model performance across a recent agentic software-engineering workflow involving:
implementation follow-through
proof quality
ability to preserve scope boundaries
responsiveness to review feedback
defect and blocker resolution
judgment around when not to overclaim readiness
This was not a synthetic benchmark. It was based on real multi-step engineering work with review gates and acceptance criteria.
Findings
1. Fable Is Strong At Execution
Fable performed well when the task was:
clearly scoped
bounded in surface area
supported by explicit acceptance criteria
focused on implementation or proof repair
part of an iterative build-review loop
It was especially effective at moving work forward quickly after reviewers identified concrete blockers.
2. Fable Benefits From Tight Direction
Fable’s quality improved when instructions were specific and constraints were explicit. It performed best when the task included:
exact scope boundaries
named proof requirements
clear “do not claim” limits
explicit sequencing
narrow next-step definitions
This suggests Fable is a strong builder model, but it should not be left to infer broad product or risk posture on its own.
3. Opus Remains Better Suited For Judgment Gates
Opus is still the safer choice for work that requires deeper reasoning, including:
final acceptance decisions
architecture tradeoffs
legal or policy-sensitive framing
product-risk analysis
public-readiness claims
roadmap ordering
adversarial review
Where mistakes are expensive, Opus remains the better default.
4. Fable Did Not Eliminate Review Burden
Fable’s recent work still required review corrections around:
proof completeness
scope discipline
over-broad status claims
evidence quality
when to continue autonomous work versus wait
These were manageable issues, but they show that Fable should not be treated as a replacement for strong independent review.
Recommended Model Strategy
Use Fable For
implementation
small-to-medium feature slices
bug fixes
test updates
documentation drafts
proof repair
structured no-code prep
fast iteration after review comments
Use Opus For
final review
high-risk correctness checks
architecture decisions
roadmap sequencing
launch-readiness decisions
public-claim approval
legal, financial, or regulated-domain judgment
adversarial analysis
Recommended Workflow
Give Fable tightly scoped implementation tasks.
Require explicit proof and acceptance artifacts.
Use Opus as an independent reviewer before accepting high-risk work.
Keep public claims gated until review confirms the exact scope.
Prefer small slices over broad autonomous mandates.
Bottom Line
Fable appears to be a productive and efficient builder model. It can increase throughput when paired with strong review discipline. Opus remains the better choice for judgment-heavy, high-risk, or externally visible decisions.
The strongest setup is a two-model workflow: Fable builds; Opus reviews and certifies.

