AI 선택의 기준이 바뀌고 있어요
법률 AI가 선도 클로즈드 모델과 맞먹는 성능을 냈는데, 비용이 런당 10분의 1이에요. 컴퓨터 사용 에이전트가 OSWorld-Verified에서 76% 이상을 기록했는데, 쓴 모델은 소형이에요. 추론 비용이 클로즈드 모델 대비 약 20배 낮은데, 에이전트 정확도는 오픈 모델 최상위예요.
NVIDIA가 7월 14일 Nemotron Labs 블로그 시리즈를 시작하면서 이 사례들을 공개했어요. 6개 기업이 Nemotron 오픈 모델을 도메인에 맞게 커스터마이징해서 달성한 결과예요. 각 기업이 직접 발표한 수치라 NVIDIA 입장의 큐레이션이라는 점은 감안해야 하지만, 이 정도 구체성이 나왔다는 건 신호예요.
NVIDIA의 핵심 주장은 이거예요: “AI에서의 경쟁 우위는 어떤 모델을 고르느냐보다 어떻게 쌓느냐에서 점점 더 많이 온다.” (원문 인용)
핵심 요약
- Harvey: Nemotron 3 Ultra 기반 법률 AI, 복잡한 법률 태스크에서 선도 클로즈드 모델 수준 달성 · 런당 비용 10배 이상 절감 (Harvey 자체 주장, 비교 모델 미특정)
- Arcee AI: 출력 100만 토큰당 약 90센트 · 클로즈드 모델 대비 약 20배 저렴 (Arcee AI 자체 주장, 비교 모델 미특정)
- H Company: Holotron 3 Nano(Nemotron 3 Nano Omni 포스트 트레이닝), OSWorld-Verified 76% 이상 (H Company 발표)
- LangChain: 모델 재훈련 없이 프롬프트·도구·미들웨어만 조정 → 오픈 모델 중 상위 에이전트 정확도, 클로즈드 모델 대비 약 10배 저렴 (LangChain 자체 주장)
- Glean: Nemotron과 대형 클로즈드 모델 페어링 → 낮은 레이턴시·적은 토큰으로 엔터프라이즈 검색 제공
- Abridge·Heidi Health·YTL AI Labs: 임상 문서화·말레이어 현지화 등 도메인 특화 사례 추가 공개
The Selection Criteria Are Shifting
Legal AI matching leading closed models — at 10x lower cost per run. A computer-use agent scoring above 76% on OSWorld-Verified with a small model. Inference costs roughly 20x below closed model equivalents, with top agent accuracy among open models.
NVIDIA launched the Nemotron Labs blog series on July 14, publishing six company cases built on Nemotron open models. These are self-reported figures curated by NVIDIA, so the framing matters — but the specificity is a signal worth tracking.
NVIDIA’s headline claim: “Increasingly, competitive AI advantage comes from how organizations build with available models, more than which one they choose.”
TL;DR
- Harvey: Nemotron 3 Ultra-based legal AI matches leading closed models on complex legal tasks at at least 10x lower cost per run (Harvey’s own claim; comparison model not specified)
- Arcee AI: roughly $0.90 per million output tokens — approximately 20x cheaper than closed models (Arcee AI’s own claim; comparison baseline not named)
- H Company: Holotron 3 Nano (post-trained Nemotron 3 Nano Omni) scores above 76% on OSWorld-Verified (H Company’s announcement)
- LangChain: No model retraining — adjusted prompts, tools, and middleware only → top agent accuracy among open models at ~10x lower cost than closed alternatives (LangChain’s own claim)
- Glean: Pairs Nemotron with larger closed models for enterprise search at lower latency with fewer tokens
- Abridge, Heidi Health, YTL AI Labs: Domain-specific cases in clinical documentation and Malay language customization
사례별로 뜯어보면
Harvey — 법률 AI에서 10배 비용 차이
법률 AI Harvey는 Nemotron 3 Ultra를 법률 도메인에 맞게 구성해서 복잡한 법률 태스크에서 “선도 클로즈드 모델과 맞먹는 성능을 런당 10배 이상 낮은 비용으로” 달성했다고 발표했어요. (Harvey 자체 주장; 비교 대상 모델 이름은 공개하지 않았어요)
법률 문서 처리는 반복 볼륨이 크고 단순 실수가 직접적인 리스크가 돼요. 이런 태스크에서 모델 비용은 운영비 그 자체예요. 비교 모델이 특정되지 않아서 수치를 그대로 받아들이기엔 한계가 있지만, 실제로 운영 중인 법률 AI가 이 방향으로 가고 있다는 건 확인된 사실이에요.
H Company — 소형 모델로 컴퓨터 사용 76%
H Company는 Nemotron 3 Nano Omni를 자체 컴퓨터 사용 데이터로 포스트 트레이닝해서 Holotron 3 Nano를 만들었어요. OSWorld-Verified에서 76% 이상의 정확도를 기록했고요. OSWorld-Verified는 파일 열기·앱 탐색·화면 기반 추론 같은 실제 컴퓨터 태스크를 평가하는 벤치마크예요.
포인트는 Nano급 — 즉 소형 모델 — 으로 이 수준을 냈다는 거예요. 컴퓨터 사용 에이전트를 항상 실행 상태로 두려면 서빙 비용이 관건이에요. 소형 모델이 이 벤치마크에서 경쟁력을 갖췄다는 건 에이전트 운영 비용 측면에서 의미 있어요.
Arcee AI — 100만 토큰당 90센트
Arcee AI는 출력 100만 토큰당 약 90센트를 달성했다고 발표했어요. 클로즈드 모델 대비 약 20배 저렴한 수준이라고 해요. (Arcee AI 자체 주장; 비교 기준 모델 미특정) 어떤 클로즈드 모델과 비교한 건지는 명시되지 않아서 배수 수치는 참고용으로 봐야 해요.
다만 90센트/1M 출력 토큰이라는 절대 수치는 확인 가능한 가격이에요. 에이전트 루프를 많이 돌리는 워크플로우라면 이 차이가 실질적인 비용 차이로 이어져요.
LangChain — 재훈련 없이 성능 올리기
이 케이스가 가장 흥미로워요. LangChain은 Nemotron 3 Ultra용 Deep Agents harness를 조정했는데, 모델을 재훈련하거나 파인튜닝하지 않고 프롬프트·도구·미들웨어만 바꿨어요. 그러면서 오픈 모델 중 상위 에이전트 정확도를 클로즈드 모델 대비 약 10분의 1 비용으로 달성했다고 해요. (LangChain 자체 주장)
“오픈 모델 커스터마이징 = 파인튜닝”이라는 생각이 있는데, 이 케이스는 그 등호가 아니라는 걸 보여줘요. 어떻게 프롬프트를 짜고, 어떤 도구를 붙이고, 미들웨어 레이어를 어떻게 구성하느냐가 성능에 직접 영향을 줄 수 있어요. 이게 가능하다면 진입 장벽이 낮아지는 거예요.
그 외 사례들
- Glean: “Waldo”라는 에이전틱 검색 모델을 만들었어요. Nemotron을 대형 클로즈드 모델과 함께 쓰는 구성으로, 낮은 레이턴시와 적은 토큰으로 엔터프라이즈 검색을 제공해요. 오픈/클로즈드 혼용 아키텍처예요
- Abridge: Nemotron 커스터마이징으로 임상 대화 특화 foundation model을 만들었어요. 의사-환자 대화 처리에 맞춘 모델이에요
- Heidi Health: frontier-scale 컴퓨트 없이 임상 문서화에서 frontier급 결과를 냈다고 해요. (구체 수치는 공개하지 않았어요)
- YTL AI Labs: Nemotron을 말레이어로 포스트 트레이닝해서 말레이시아 개발자 커뮤니티용 현지화 AI를 배포했어요
Nemotron Labs와 Nemotron Coalition이 뭔가요
Nemotron Labs는 제품이나 프로그램이 아니에요. NVIDIA가 이번에 시작한 블로그 시리즈 이름이에요. “오픈 모델·데이터셋·훈련 기법으로 기업이 특화 AI를 만드는 방법”을 다루는 시리즈예요.
함께 언급된 Nemotron Coalition은 모델 빌더와 개발자가 데이터·평가·도메인 지식을 공유해서 Nemotron을 개선하는 생태계 구상이에요. 해커톤 제출물이나 커뮤니티 기여를 통해 산업별 재사용 가능한 자산을 만드는 방향이라고 해요. (창립 멤버 목록과 설립일은 발표문에 명시되지 않았어요)
모델에 접근하려면 build.nvidia.com에서 가능해요.
왜 중요한가요
오픈 vs 클로즈드 논쟁은 오래됐지만, 이번 발표가 다른 건 구체적인 회사 이름과 수치가 붙기 시작했다는 거예요.
“오픈 모델로도 된다”는 말은 많았어요. “Harvey가 법률 태스크에서 클로즈드 모델과 맞먹고 비용은 10분의 1”이라는 말은 다른 무게예요. 수치가 모두 기업 자체 발표라 독립 검증이 필요하지만, 이 정도 구체성이 동시에 여러 도메인에서 나왔다는 건 신호예요.
AI 도입을 검토하는 기업 입장에서 실질적인 변수가 하나 더 추가됐어요. 클로즈드 모델 API 하나로 가는 게 빠르고 편해요. 그런데 도메인이 명확하고 볼륨이 크다면, 오픈 모델 커스터마이징이 단순히 더 저렴한 게 아니라 전략적으로 다른 옵션이에요 — 외부 의존도를 줄이고, 도메인 지식을 직접 모델에 쌓고, 데이터 주권을 가져갈 수 있거든요.
반대로 말하면, 볼륨이 작거나 도메인이 일반적이거나 빠른 실험이 목적이라면 클로즈드 API가 여전히 합리적이에요. 이건 기술 선택의 문제가 아니라 비즈니스 맥락에 따른 판단이에요.
핵심 통찰
AI 경쟁력이 ‘어떤 모델 API를 쓸까’에서 ‘어떻게 쌓을까’로 무게 중심을 옮기고 있어요.
클로즈드 모델이 나쁜 게 아니에요. 빠르고 강력하고 시작하기 쉬워요. 다만 도메인 특화·비용 민감·데이터 주권이 모두 중요한 영역에서는 오픈 모델 커스터마이징이 이제 이론이 아니라 실제 배포 사례로 존재해요. NVIDIA가 이 흐름을 생태계 서사로 포장하는 건 NVIDIA 입장의 이야기지만, 6개 사례의 내용 자체는 살펴볼 만해요.
My Take
이 발표를 액면 그대로 받아들이긴 어려워요. 모든 수치는 기업 자체 주장이고, 비교 대상 클로즈드 모델이 대부분 이름을 밝히지 않았어요. NVIDIA가 성공 사례만 큐레이팅했다는 것도 당연한 거고요.
그럼에도 LangChain 케이스는 실무적으로 따져볼 만해요. 모델 재훈련 없이 프롬프트·도구·미들웨어 레이어 조정만으로 오픈 모델의 에이전트 성능을 끌어올릴 수 있다면, 진입 비용이 생각보다 낮아요. 파인튜닝 인프라가 없어도 시작할 수 있다는 얘기예요.
현재 클로즈드 모델 API를 쓰면서 비용이 부담되거나, 특정 도메인에서 품질이 아쉽다면, 오픈 모델 커스터마이징을 검토해볼 시점이 왔어요. Harvey와 Arcee AI의 기술 블로그가 출발점으로 좋아요.
Breaking Down the Cases
Harvey — 10x Lower Cost in Legal AI
Legal AI company Harvey built a system on Nemotron 3 Ultra that matches “leading closed models on complex legal tasks at at least 10x lower cost per run.” (Harvey’s own claim; no specific comparison model named)
Legal document processing runs at high volume, and errors carry direct liability risk. In that context, model cost is operating cost. The comparison baseline isn’t named, so the 10x figure needs independent verification — but Harvey is in production, and the direction is documented.
H Company — 76%+ on Computer Use With a Small Model
H Company post-trained Nemotron 3 Nano Omni on its proprietary computer-use data to build Holotron 3 Nano. It scored above 76% accuracy on OSWorld-Verified — a benchmark covering real computer tasks like opening files, navigating applications, and screen-based reasoning. (H Company’s announcement)
The key point: a Nano-class (small) model achieving this. Running a computer-use agent continuously makes serving cost a practical constraint. A small model that’s competitive on this benchmark changes the economics of always-on agentic workflows.
Arcee AI — 90 Cents Per Million Output Tokens
Arcee AI reports reaching roughly $0.90 per million output tokens — approximately 20x cheaper than closed models. (Arcee AI’s own claim; the comparison baseline isn’t named, so the multiplier is hard to verify.)
The absolute number — $0.90/M output tokens — is a verifiable price point. For workloads running agentic loops at scale, that cost differential translates directly to operating expense.
LangChain — Better Agents Without Retraining
This is the most interesting case. LangChain tuned its Deep Agents harness for Nemotron 3 Ultra by adjusting prompts, tools, and middleware — no model retraining or fine-tuning required. Result: top agent accuracy among open models at roughly 10x lower cost than closed alternatives. (LangChain’s own claim)
“Open model customization” often implies fine-tuning. This case shows that’s not the only path. Getting the prompt design, tooling, and middleware layer right can drive substantial performance gains. The implication for teams without training infrastructure: the barrier to entry may be lower than expected.
Other Cases
- Glean: Built “Waldo” — pairs Nemotron with larger closed models for enterprise search at lower latency with fewer tokens. A hybrid open/closed architecture
- Abridge: Used Nemotron customization to build what it calls “the first foundation model purpose-built for clinical conversations”
- Heidi Health: Claims frontier-quality clinical documentation outcomes without frontier-scale compute (no specific metrics published)
- YTL AI Labs: Post-trained Nemotron in Malay to build localized AI for Malaysia’s developer community
What Is Nemotron Labs?
Nemotron Labs isn’t a product or program. It’s the name of the blog series NVIDIA launched — documenting “how the latest open models, datasets and training techniques help businesses build specialized AI systems.”
Also introduced: the Nemotron Coalition, described as “helping turn open model development into an ecosystem effort, bringing model builders and developers together to improve Nemotron through shared data, evaluations and domain expertise.” (Founding member list and launch date not specified in the announcement) Access to models is available through build.nvidia.com.
Why It Matters
The open vs. closed model debate has been ongoing. What’s different now: company names and specific numbers are showing up alongside the claims.
“Open models can work” is a general assertion. “Harvey matched closed models on legal tasks at 10x lower cost” is a different category of statement. All figures are self-reported and unverified against named baselines — but multiple companies across multiple domains publishing cases at this level of specificity is a signal worth noting.
For companies evaluating AI deployment, there’s now a clearer question on the table. Closed model APIs are fast to start with and powerful. But if your domain is well-defined and your volume is high, open model customization isn’t just cheaper — it’s strategically different. Lower external dependency, domain knowledge baked directly into the model, and data sovereignty.
Conversely, for low-volume or general-purpose use cases, or when speed of experimentation matters most, closed APIs remain the practical choice. This is a business context question, not a technology preference.
Key Insight
The center of gravity for building competitive AI is shifting from “which model API to call” to “how to build on top of available models.”
Closed models aren’t going away — they’re fast, powerful, and easy to start with. But for domain-specific, cost-sensitive, or data-sovereignty use cases, open model customization now exists as a documented production option, not just a theoretical alternative. NVIDIA is framing this as its ecosystem narrative — naturally — but the six cases are worth reading on their own terms.
My Take
Take the figures with appropriate skepticism. Every number here is a company self-report, and most comparison baselines are unnamed. This is NVIDIA curating its best case studies.
The LangChain case is worth examining closely from a practical standpoint: if prompt, tool, and middleware adjustments — without fine-tuning — can drive this kind of performance improvement on top of open models, the barrier to entry is lower than most assume. You don’t need training infrastructure to start.
If you’re running workloads on closed model APIs where cost is a constraint or domain quality is lacking, this is a reasonable time to evaluate open model alternatives. Harvey’s and Arcee AI’s technical write-ups are good starting points.
댓글Comments