6.5년치 보안 점검을 20시간에 — 앨버타 주정부가 Claude로 한 일6.5 Years of Security Work in 20 Hours: How Alberta's Government Used Claude

6.5 Years of Security Work in 20 Hours: How Alberta's Government Used Claude

캐나다 앨버타 주정부가 Claude Code와 약 50개 자율 에이전트를 병렬 운용해 466,000,000줄의 코드를 20시간 만에 보안 점검했어요. 25년 된 레거시 포털을 4–5일 만에 재구축한 사례도 포함됩니다.Alberta's government used Claude Code and about 50 autonomous agents to security-scan 466 million lines of code in 20 hours — work that would have taken 6.5 years the traditional way.

Source

Government of Alberta uses Claude to find and fix cybersecurity vulnerabilities across government systems

전통 방법으론 6.5년, Claude로는 20시간이었어요

캐나다 앨버타 주정부 기술혁신부(Ministry of Technology and Innovation)가 Claude Code로 27개 부처 전체 IT 시스템을 보안 점검했어요. 466,000,000줄의 코드, 1,280개 애플리케이션, 3,400개 코드 리포지토리. 전통적인 방법으로는 6.5년이 걸릴 분량을 20시간 안에 처리한 거예요.

이 사례가 흥미로운 건 규모 때문만은 아니에요. 어떻게 구성했는지가 더 중요한 이야기예요.

핵심 요약

  • 466,000,000줄의 코드를 20시간 만에 보안 점검. 전통 방식 대비 약 6.5년 단축.
  • 약 50개의 자율 에이전트(Claude Agent SDK)를 병렬 운용. 레드팀·블루팀·품질 검토 에이전트로 역할 분담.
  • 애플리케이션당 약 95개 보안 통제 항목 점검.
  • 25년 된 레거시 Java 포털(원래 구축에 5개월 소요)을 Claude Code로 4–5일 만에 재구축.
  • 2026년 가을부터 전 주정부 시스템으로 프로그램 확대 예정.

20 Hours to Do What Traditionally Takes 6.5 Years

Alberta’s Ministry of Technology and Innovation used Claude Code to security-scan the entire provincial IT infrastructure across 27 ministries. The scope: 466,000,000 lines of code, 1,280 applications, and 3,400 code repositories — accomplished in under 20 hours. The traditional approach would have taken an estimated 6.5 years.

The scale is remarkable. But the architecture of how they did it matters more.

TL;DR

  • 466 million lines of code security-scanned in 20 hours. Traditional approach: approx. 6.5 years.
  • About 50 autonomous agents deployed in parallel using the Claude Agent SDK — red team, blue team, and quality review agents working together.
  • About 95 security controls checked per application.
  • A 25-year-old legacy Java portal (originally 5 months to build) rebuilt in 4–5 days.
  • Program expanding across all government systems starting fall 2026.

에이전트를 어떻게 배치했나

두 단계 스캐닝

보안 점검은 두 단계로 설계됐어요. 먼저 규칙 기반 플래깅(rule-based flagging) 단계에서 의심 패턴을 걸러내고, 두 번째 단계에서 에이전트가 구체적인 파일명과 줄 번호를 제시하며 상세 검토해요. “취약점이 있을 수 있다”는 판정에 그치지 않고, 수정 코드와 테스트까지 함께 생성하는 방식이에요.

약 50개의 자율 에이전트(autonomous agent)가 동시에 병렬 실행됐어요. Claude Agent SDK로 구성됐고, 역할은 세 가지로 나뉘어요:

  • 레드팀 에이전트: 외부 공격자의 관점에서 취약점을 탐색
  • 블루팀 에이전트: 국제 보안 표준 대비 방어력을 평가
  • 품질·명확성 검토 에이전트: 산출물의 정확도와 재현 가능성을 검사

각 애플리케이션당 약 95개의 보안 통제 항목을 점검해요.

레거시 코드의 재구축

25년 된 레거시 Java 보조금 포털이 이 프로그램의 대표적인 사례예요. 원래 구축에 5개월이 걸렸던 시스템을 4–5일 만에 재구축했어요. 한 부처는 185개 레거시 애플리케이션을 16개 현대 애플리케이션으로 통합할 계획을 세우고 있기도 해요.

앨버타 주의 IT 시스템에는 세금 기록, 조달 정보, 사회서비스 케이스 파일 같은 고민감 데이터가 들어 있어요. 기술적 부채(technical debt) 규모는 수십억 달러로 추산돼요(주정부 추정치). 이 점검 프로그램은 2026년 가을부터 전 주정부 시스템으로 확대될 예정이에요.

사람 교육도 함께

기술 도입과 병행해서 앨버타 AI Academy를 운영하고 있어요. 수천 명의 공무원이 AI 기초 과정을 이수했고, 10,000명 이상의 일반 시민이 기업 애플리케이션 교육에 참여했어요.

왜 중요한가

“AI가 코드를 쓴다”는 이야기는 이제 익숙해요. 그런데 이 사례는 한 발 더 나간 거예요. 에이전트들이 서로 협력해 보안 업무 전체 사이클(탐색→검증→수정→테스트)을 에이전트들이 자율적으로 처리했어요.

공공 섹터에서 이 규모로 실증된 사례는 드물어요. 민간 기업 대비 데이터 민감도와 레거시 시스템 복잡도가 훨씬 높은 환경이었다는 점에서, 기업 IT 부서가 직접 참고할 수 있는 레퍼런스예요.

비용 측면에서는 6.5년의 인건비와 20시간의 GPU 비용을 굳이 비교할 필요가 없을 정도예요. 단, 에이전트 설계와 프롬프트 엔지니어링에 얼마나 시간이 들었는지는 이 발표에서 구체적으로 공개하지 않았어요.

핵심 통찰

에이전트 AI의 ROI는 “한 작업을 얼마나 빠르게”가 아니라 “병렬로 몇 개의 작업을 동시에”에서 나온다.

50개 에이전트가 병렬로 달리는 구조가 6.5년 → 20시간을 만든 본질이에요. 단일 에이전트가 아무리 빠르게 일해도 이 배율은 나오지 않아요. AI 도입을 설계하는 질문이 “무슨 모델을 쓸까”에서 “몇 개의 에이전트를 어떻게 역할 분담할까”로 옮겨가고 있어요.

My Take

이 사례에서 두 가지가 눈에 들어왔어요.

하나는 레거시 해소의 경제학이에요. 25년 된 Java 포털을 5일에 재구축했다는 건, 그동안 “손대기 너무 어렵다”는 이유로 미뤄온 기술적 부채 상환의 비용 구조가 바뀌었다는 신호거든요. 레거시 마이그레이션 프로젝트의 ROI 계산을 처음부터 다시 해봐야 하는 시점일 수 있어요.

다른 하나는 보안 에이전트의 자기 검증 구조예요. 레드팀이 공격하고 블루팀이 방어를 평가하는 구조는 단순히 빠른 게 아니라 교차 검증을 내장하고 있어요. 이 원리는 보안 외 영역에도 바로 적용할 수 있어요. 초안 에이전트 + 비평 에이전트 + 팩트체크 에이전트가 협력하는 식으로요.

직접 해볼 수 있는 시작점이 하나 있어요. Claude Agent SDK의 multi-agent 기능을 열고, 가장 반복적인 내부 검토 프로세스 하나를 에이전트 2–3개로 분해해보는 거예요. 50개까지 갈 필요 없이, 2개만 병렬로 돌아도 체감은 달라져요.

How the Agents Were Deployed

Two-Phase Scanning

The security review ran in two stages. First, rule-based flagging filtered suspicious patterns. Then agents conducted detailed analysis — pointing to specific file names and line numbers, generating remediation code, writing tests, and rewriting legacy code where needed. The output was actionable, not just diagnostic.

About 50 autonomous agents ran in parallel, built on the Claude Agent SDK, with three distinct roles:

  • Red team agents: searching for vulnerabilities from an attacker’s perspective
  • Blue team agents: evaluating defenses against international security standards
  • Quality and clarity review agents: verifying accuracy and reproducibility of outputs

Each application was checked against roughly 95 security controls.

Legacy Code, Rebuilt

A 25-year-old legacy Java benefits portal — originally a 5-month build — was rebuilt in 4–5 days. One ministry is planning to consolidate 185 legacy applications into 16 modern ones.

Alberta’s systems contain highly sensitive data: tax records, procurement details, social services case files. The province estimates its technical debt runs into the billions (government estimate). The program expands province-wide in fall 2026.

Training People Too

Alongside the technical work, Alberta’s AI Academy has trained thousands of civil servants on AI fundamentals. Over 10,000 citizens have participated in enterprise application courses.

Why It Matters

“AI writes code” is familiar. This case goes further: agents collaborated across the full security cycle — detection, verification, remediation, testing — autonomously.

Large-scale government deployments like this are rare. The combination of high data sensitivity, deeply entangled legacy systems, and strict compliance requirements makes the public sector a harder environment than most private companies. That this worked here makes it a meaningful reference for enterprise IT teams.

On cost: comparing 6.5 years of labor to 20 hours of GPU compute doesn’t require a calculator. What isn’t disclosed: how much time went into agent design and prompt engineering. That’s the honest caveat.

Key Insight

The ROI of agentic AI comes not from “how fast can one agent work” but from “how many agents can work in parallel.”

50 agents running concurrently produced the 6.5 years → 20 hours transformation. A single agent running faster wouldn’t get there. The design question for AI adoption is shifting from “which model?” to “how many agents, with what role configuration?”

My Take

Two things stand out.

First, the economics of legacy resolution. A 5-month build, rebuilt in 5 days. Projects deferred for years as “too complex to touch” may now be viable again. The ROI math for legacy migration deserves a fresh look.

Second, the self-validating structure of adversarial agents. Red team attacks, blue team defends and evaluates. That’s not just fast — it’s cross-verified by design. The pattern applies beyond security: draft agent + critique agent + fact-check agent, running in a loop.

A starting point: open the Claude Agent SDK’s multi-agent features and decompose one of your most repetitive internal review processes into 2–3 agents. You don’t need 50. Two agents running in parallel already changes the feel.

댓글Comments