성능을 살 건지, 비용 효율을 살 건지 — Anthropic이 기준을 세웠어요
지난 몇 년간 AI 모델 경쟁은 “누가 더 강한가”로만 흘렀어요. 그런데 Anthropic이 2026년 7월 24일 내놓은 Claude Opus 5는 다른 질문을 던져요. “같은 결과를 얼마나 싸게 낼 수 있는가.” 프런티어 모델 Claude Fable 5의 절반 가격으로 그에 근접한 성능을 낸다는 포지셔닝이에요.
가격은 입력 $5/백만 토큰, 출력 $25/백만 토큰 — Opus 4.8과 같고, Fable 5의 절반 수준이에요.
핵심 요약
- Anthropic이 2026-07-24 Claude Opus 5를 출시했어요. Claude.ai·Claude Code·API·Claude Cowork에서 즉시 이용 가능해요
- 가격은 Opus 4.8과 동일하고($5/$25 per M 토큰), Fable 5 대비 약 절반이에요
- 에이전트 코딩(CursorBench 3.2), 자동화(Zapier AutomationBench), 컴퓨터 사용(OSWorld 2.0), 과학 연구 작업에서 경쟁 모델 대비 상위권 결과를 냈어요
- 정렬(alignment) 자동화 감사에서 역대 최저 불일치 점수 2.3을 달성했어요
- Fast 모드(기본 속도의 2.5배, 가격 2배)와 중간 대화 중 도구 변경(베타) 기능이 새로 추가됐어요
The line Anthropic just drew: performance vs. cost-efficiency
For years, AI model competition ran on a single axis — who’s more capable? Claude Opus 5, released on July 24, 2026, shifts the framing. Anthropic positions it as a “thoughtful and proactive model, offering near-frontier-level intelligence at roughly half the cost of Claude Fable 5.” The question isn’t whether it tops every chart. It’s whether it delivers enough, cheaper.
Pricing: $5 input / $25 output per million tokens — unchanged from Opus 4.8, and roughly half of Fable 5.
TL;DR
- Anthropic released Claude Opus 5 on 2026-07-24, available immediately on Claude.ai, Claude Code, API, and Claude Cowork
- Priced identically to Opus 4.8 ($5/$25 per M tokens), roughly half of Fable 5
- Top-tier results on agent coding (CursorBench 3.2), automation (Zapier AutomationBench), computer use (OSWorld 2.0), and scientific research tasks
- Lowest misalignment score on record in automated behavioral audits: 2.3
- New: Fast mode (2.5× speed, 2× price) and mid-conversation tool changes (beta)
Claude Opus 5, 어떤 모델인가요
포지션: 프런티어 바로 아래, 에이전트 실무용
Anthropic은 Opus 5를 “thoughtful and proactive model”로 소개했어요. 최강 모델을 쓰기엔 비용이 부담스럽고, 가벼운 모델로는 복잡한 작업이 버거운 팀에게 맞는 자리예요. 에이전트 코딩·장기 실행 작업·지식 집약 업무를 주된 용도로 설계했어요.
Fast 모드를 켜면 기본 속도의 2.5배로 동작하고, 가격은 기본의 2배가 돼요. 추론 속도가 중요한 에이전트 워크플로우에서 트레이드오프를 직접 선택할 수 있는 옵션이에요.
벤치마크: 어디서 두드러지나요
1차 소스(anthropic.com/news)에서 확인된 상대 성능이에요. 절대 점수는 System Card를 참고해야 해요.
| 벤치마크 | Opus 5 결과 |
|---|---|
| Frontier-Bench v0.1 | Opus 4.8 대비 2배 이상, 전 경쟁사 초과 |
| CursorBench 3.2 | Fable 5 최고점 0.5% 범위 내, 절반 비용 |
| ARC-AGI 3 | 차선 모델 대비 3배 |
| Zapier AutomationBench | 차선 모델 대비 약 1.5배 통과율 |
| OSWorld 2.0 | Fable 5 최고점을 1/3 비용으로 초과 |
과학 연구 분야에서도 Opus 4.8 대비 유기화학 작업 10.2포인트, 단백질 관련 작업 7.7포인트 향상이 확인됐어요.
파트너사 코멘트도 같은 방향을 가리켜요. Cursor의 공동창업자 Sualeh Asif는 “Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost”, Zapier의 CEO Wade Foster는 “Claude Opus 5 topped Zapier’s AutomationBench leaderboard without spending more tokens”라고 했어요.
안전성: 정렬 점수를 처음으로 2점대로
안전성 지표에서 주목할 수치는 자동화 행동 감사(automated behavioral audits)의 불일치 점수 2.3이에요. Anthropic이 공개한 최저 수치이고, Opus 4.8·Sonnet 5·Fable 5 대비 가장 우수한 결과예요.
사이버보안 쪽에서 분류기 개입 빈도는 Fable 5 대비 약 85% 감소할 것으로 예상돼요. 소스 코드 취약점 식별은 허용하고, 바이너리 기반 취약점 스캔·침투 테스트·악용 코드 생성은 차단하는 구조예요.
Cyber Verification Program(CVP)도 함께 공개됐어요. 보안 연구 목적으로 쓸 때 거치는 공식 검증 절차예요.
새 기능: 에이전트 개발자를 위한 베타 2종
- 중간 대화 도구 변경: 긴 에이전트 세션 중 프롬프트 캐시 무효화 없이 도구를 바꿀 수 있어요. 기존에는 도구를 변경하면 캐시가 날아가서 비용이 다시 불어났는데, 이 문제를 없앴어요
- 자동 폴백: 안전 분류기가 요청에 플래그를 달면 다른 모델로 자동 라우팅해요. 에이전트 파이프라인에서 오류 처리를 별도로 짜지 않아도 돼요
왜 중요한가요
가격 경쟁의 시작점이 프런티어 모델 바로 아래까지 올라왔어요.
기존 구도는 “저가 모델 vs 비싼 프런티어 모델”이었어요. Opus 5는 그 틈을 파고들어, “프런티어에 가까운 성능을 프런티어 가격의 절반에 쓸 수 있다”는 새로운 기준을 제시했어요. 에이전트를 대규모로 운영하는 팀에게는 이 기준이 모델 선택의 실질적인 변수가 돼요.
벤치마크 구성도 달라졌어요. Frontier-Bench, CursorBench, AutomationBench, OSWorld — 모두 단순 지식 테스트가 아니라 실제 작업 수행 능력을 재는 지표예요. AI 모델 평가 기준이 “얼마나 아는가”에서 “얼마나 하는가”로 옮겨가고 있다는 신호예요.
핵심 통찰
에이전트 시대의 경쟁 기준은 “최고 성능”이 아니라 “단위 작업당 비용”으로 이동하고 있어요.
이 흐름에서 중요한 건 절대 성능이 아니에요. Opus 5가 Fable 5를 모든 벤치마크에서 이기지 않아도, “절반 가격에 0.5% 차이 안”이라면 에이전트를 운영하는 팀에게는 Opus 5가 합리적인 기본 선택이 돼요. Anthropic이 이 기준을 먼저 만들었어요.
My Take
Opus 5 출시에서 가장 주목한 건 벤치마크 선택이에요. ARC-AGI 3, OSWorld 2.0 — 단순 추론 테스트가 아니에요. 컴퓨터를 실제로 다루고, 연구 과제를 수행하고, 자동화 흐름을 통과하는 능력을 재는 지표예요. Anthropic이 이 지표들로 성능을 정의했다는 건, 앞으로의 모델 경쟁이 “chat 품질”이 아니라 “에이전트 작업 완료율”로 흘러갈 수도 있다는 방향을 시사해요.
에이전트를 직접 쓰거나 만드는 분이라면 지금 바로 해볼 수 있는 게 있어요. Claude API에서 Opus 4.8 대신 Opus 5로 교체해보고, 품질이 유지되는지 확인하는 거예요. Fast 모드도 병렬 에이전트 실행에서 어떻게 달라지는지 한 번 돌려볼 가치가 있어요. 비용이 같거나 비슷하다면, 안 바꿀 이유가 없어요.
What Claude Opus 5 actually is
Positioning: sub-frontier, built for agent workflows
Anthropic describes Opus 5 as a “thoughtful and proactive model” — designed for the space between lightweight models that can’t handle complex tasks and frontier models that cost too much to run at scale. Primary use cases: agent coding, long-horizon tasks, knowledge-intensive work.
Fast mode offers 2.5× the base speed at 2× the price — a deliberate tradeoff option for latency-sensitive agent pipelines.
Benchmarks: where it stands out
Relative performance confirmed from the primary source (anthropic.com/news). Absolute scores are in the System Card.
| Benchmark | Opus 5 Result |
|---|---|
| Frontier-Bench v0.1 | More than doubles Opus 4.8; beats all competitors |
| CursorBench 3.2 | Within 0.5% of Fable 5’s peak at half the cost |
| ARC-AGI 3 | 3× next-best model |
| Zapier AutomationBench | ~1.5× next-best model pass rate |
| OSWorld 2.0 | Surpasses Fable 5’s best result at one-third the cost |
On scientific research: +10.2 points on organic chemistry tasks, +7.7 points on protein-related tasks (both vs. Opus 4.8).
Partner feedback points in the same direction. Cursor co-founder Sualeh Asif: “Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost.” Zapier CEO Wade Foster: “Claude Opus 5 topped Zapier’s AutomationBench leaderboard without spending more tokens.”
Safety: misalignment score hits 2.3
The key safety figure is a misalignment score of 2.3 in automated behavioral audits — Anthropic’s lowest on record, and better than Opus 4.8, Sonnet 5, and Fable 5.
On cybersecurity: classifier intervention frequency is expected to drop approximately 85% vs. Fable 5. Source code vulnerability identification is allowed; binary-based scanning, penetration testing, and exploit generation are blocked.
The Cyber Verification Program (CVP) launches alongside — a formal verification pathway for security research use cases.
New features: two betas for agent developers
- Mid-conversation tool changes: Swap tools during a long agent session without invalidating the prompt cache. Previously, changing tools mid-session would flush the cache and reset costs. This removes that penalty.
- Automatic fallback: If the safety classifier flags a request, it auto-routes to another model instead of returning an error. Agent pipelines no longer need custom error handling for safety-triggered failures.
Why it matters
The price competition just reached the near-frontier tier.
The old dynamic was cheap models vs. expensive frontier models. Opus 5 steps into the gap: near-frontier performance at half the frontier price. For teams running agents at scale, this changes the model selection calculus — not just a benchmark footnote.
The benchmark choices also signal a broader shift. Frontier-Bench, CursorBench, AutomationBench, OSWorld — none of these are knowledge quizzes. They measure actual task completion. The industry’s evaluation axis may be shifting from “what does it know” to “what can it do.”
Key Insight
In the agent era, the competitive metric isn’t peak performance — it’s cost per completed task.
Opus 5 doesn’t need to beat Fable 5 on every chart to win. “Within 0.5% at half the cost” is a rational default choice for anyone running agents at volume. Anthropic set this standard first.
My Take
The benchmark selection is what I’m watching closely. ARC-AGI 3, OSWorld 2.0 — these aren’t reasoning puzzles. They test whether a model can operate a computer, execute research tasks, and move through automation flows. Anthropic choosing these metrics to define Opus 5’s value is a signal that model competition could shift from chat quality to agent task completion rates.
If you’re building or running agents, this is worth a quick test today. Swap Opus 4.8 for Opus 5 in your Claude API setup and check whether quality holds. Try Fast mode for parallel agent runs. If the cost is the same or close, there’s no reason not to switch.
댓글Comments