Essay 006 · AI & Judgment

The Cost of Checking 確かめる費用

As AI becomes more accurate, errors decline. But fewer errors do not necessarily mean that certainty becomes cheaper. The Jagged Technological Frontier, stress testing in financial markets, and the rising cost of assurance in a world where answers are becoming abundant. AIの精度が上がれば、誤りは減る。だが、誤りが減ることと、確かさを得る費用が下がることは同じではない。Jagged Technological Frontier、金融市場のストレステスト、そして「答え」が安価になる時代の検証について。

June 2026 AI · tail risk · certainty AI · tail risk · 確かさの値段 11 min read 日本語 約11分

As AI becomes more capable, it is natural to assume that the human burden should diminish. Fewer errors should mean less checking, and eventually perhaps no human verification at all. If accuracy is imagined as a single line moving steadily upward, the conclusion seems almost unavoidable.

AIの性能が上がれば、人間の仕事は軽くなる。誤りが減れば確認の手間も減り、やがて人間による検証そのものが不要になる――精度を一本の直線として考えるなら、そう結論するのが自然である。

Reality is less obliging. When AI output was crude, its errors were often visible at a glance. It invented sources, confused numbers and stumbled halfway through an argument. Users distrusted it almost by default. They returned to the original document, recalculated the number, or reconstructed the reasoning themselves. The very immaturity of the system acted as a kind of warning label.

しかし、実際にはもう少し厄介なことが起きる。AIの出力が粗かった頃には、一読して分かる間違いが頻繁に混じっていた。存在しない資料を引用し、数字を取り違え、論理の途中で転ぶ。人は当然のように疑い、自分で原典を探し、数字を引き直した。未熟さが露呈していたからこそ、使う側にも警戒心が残っていたのである。

Once nine answers out of ten, or ninety-nine out of a hundred, become fluent and correct, the remaining error changes character. It does not disappear. It sinks into a much larger body of correct output and becomes a residual error that is increasingly difficult to distinguish from everything around it.

ところが、十のうち九つ、百のうち九十九が自然で正しい答えを返すようになると、残された一つの誤りの性質が変わる。誤りは消滅するのではない。大量の正答の内部へ沈み込み、発見しにくい残存誤差になる。

Key point要点 AI makes answers cheaper. It does not necessarily make certainty cheaper. AIが安くするのは答えであって、確信ではない。

1Error moves into the tail誤りは、tailへ沈んでいく

A financial-market analogy is useful. A decline in average losses does not mean that tail risk has disappeared. Frequent, smaller accidents may become less common, while the events that remain become rarer and more difficult to observe. The cost of detecting a Black-Swan-like failure can therefore rise sharply even as average performance improves.

金融市場のリスクに例えれば、平均的な損失率が下がったからといって、tail riskが消えたわけではないのと似ている。頻繁に起きる小さな事故は減る。しかし、残された事故は稀になり、そのブラック・スワン的な事故を捕捉するための検証コストは、むしろ飛躍的に上がる。

Financial risk management has long dealt with a version of this problem. We do not rely solely on statistical measures calibrated to normal conditions; we also replay stress scenarios against the current portfolio. The Lehman shock remains part of that vocabulary, but the dictionary has continued to expand. The 2022 UK LDI crisis exposed a violent interaction between rates, leverage and collateral calls. The failure of Silicon Valley Bank in 2023 showed how interest-rate risk could combine with liquidity risk and an extraordinarily rapid run on deposits. In April 2025, the tariff shock in the United States generated simultaneous volatility across equities, Treasuries, credit and foreign exchange.

金融のリスク管理では、こうしたtailを捉えるため、通常時の統計的なリスク量とは別にストレステストを行う。リーマン・ショック時の市場変動を現在のポートフォリオへ当て直すだけではない。その後も新しい危機が起きるたび、シナリオの辞書は増えてきた。2022年の英国LDI危機では、30年物英国債利回りがわずか四日間で140bp上昇し、それまでの歴史的変動を大きく上回った。2023年のシリコンバレー銀行の破綻では、金利リスクが流動性リスクと結びつき、預金流出が従来の想定よりはるかに速い速度で進んだ。2025年4月のいわゆるトランプ関税ショックでは、米国の通商政策変更を契機に株式、国債、クレジット、為替のボラティリティが同時に上昇し、米国債市場の流動性も急速に悪化した。

Yet stress testing contains its own limitation. A historical stress scenario does no more than map a shock that has already occurred onto today's positions. Once a Black Swan has happened, it is demoted the next day into a known stress scenario. That is why serious risk management cannot stop at historical replay. It also uses hypothetical scenarios, reverse stress testing and judgments about emerging vulnerabilities.

だが、ここにはストレステストそのものの限界がある。ヒストリカル・ストレスが測っているのは、過去に一度起きた極端な変動を、現在のポジションへ写像した場合の損失にすぎない。一度起きたブラック・スワンは、翌日から既知のストレスシナリオへ格下げされる。だから実務では、過去の再現だけでなく、まだ経験していない仮想シナリオやリバース・ストレステストまで必要になる。バーゼルのストレステスト原則も、過去の事象だけでなく、現在の脆弱性や新たに立ち上がりつつあるリスクを踏まえた仮想的な将来事象を考慮するよう求めている。

Key point要点 Known tails can be measured. What is truly dangerous is the tail that does not yet have a name. 既知のtailは測れる。だが、本当に怖いのは、まだ名前のついていないtailである。

AI assurance is likely to confront the same problem. Yesterday's hallucination becomes today's benchmark. A failure mode, once identified, can be turned into a test case and engineered against. Average accuracy rises as known errors are progressively eliminated. But the closer the testing regime comes to covering what we already understand, the larger the relative importance of failures that have not yet been classified as failures at all.

Improving accuracy, in other words, does not simply reduce risk. It also changes the composition of risk: known errors recede, while unknown errors matter more.

AIの検証も、おそらく同じ場所へ向かう。昨日発見されたハルシネーションや失敗のパターンは、今日にはベンチマークやテストケースへ組み込める。既知の誤りを一つずつ潰せば、平均精度は上がる。しかし検証体系が精緻になるほど、最後に残るのは、まだ失敗のパターンとして定義されていない誤りである。精度の上昇とは、単にリスクが減ることではない。既知の誤りが減り、未知の誤りの比重が相対的に増すことでもある。

surface of apparent correctness errors break the surface — you see them errors slip below — you trust them low accuracy high accuracy → 正しさに見える水面 誤りが水面を破る — 見える 誤りが水面下へ沈む — 信じてしまう 低精度 高精度 →
Fig. 1 — Accuracy and the visibility of error. When accuracy is low, errors protrude above the surface. As the level of correct output rises, the residual errors sink beneath it. They have not vanished. You simply have to dive deeper to find them.図1 精度と、誤りの可視性。精度が低いとき、誤りは水面を破って露出している。精度が上がるにつれ、正しい出力の水位が上がり、残存する誤りは水面下へ沈んでいく。誤りがなくなったのではない。見つけるために、より深く潜らなければならなくなったのである。

2The Jagged Technological FrontierJagged Technological Frontier

There is one image that has stayed with me when thinking about this problem. This June, in Competing in the Age of AI at Harvard Business School, Marco Iansiti introduced us to the idea of the Jagged Technological Frontier — the notion that the boundary of AI capability is not a smooth curve but an irregular, serrated edge.

この問題を考えるうえで、私には忘れにくい一枚の図がある。この六月、ハーバード・ビジネス・スクールの Competing in the Age of AI でMarco Iansiti教授から学んだ、Jagged Technological Frontier――AIの能力境界が、滑らかな曲線ではなく鋸歯状に入り組んでいるという考え方である。

The term itself did not originate with Professor Iansiti. It emerged from field research by scholars at Harvard Business School and collaborators working with Boston Consulting Group, involving 758 BCG consultants. Participants were assigned to conditions without AI, with GPT-4, and with GPT-4 after receiving an overview of prompt engineering, and were asked to perform tasks designed to resemble real consulting work.

この概念は、ハーバード・ビジネス・スクールとボストン・コンサルティング・グループの研究者らが758名のBCGコンサルタントを対象に行ったフィールド実験から生まれた。参加者をAIなし、GPT-4あり、さらにプロンプト・エンジニアリングの概要を学んだうえでGPT-4を使う群に分け、実際のコンサルティングに近い複数の知識労働を行わせた。研究はワーキングペーパーとして広く知られるようになり、2026年には Organization Science に正式掲載されている。

What the study revealed was that human intuition about the order in which machines acquire competence can be misleading. AI does not simply master the easy tasks first and then proceed gradually toward the difficult ones. It may fail on something that looks elementary while performing remarkably well on a task that appears much harder. Within the same workflow, tasks on which AI materially augments human performance can sit beside tasks on which using AI makes performance worse.

That irregular boundary is the “jagged” frontier.

研究が示したのは、「AIは簡単な仕事から順番にできるようになる」という人間側の直感が、必ずしも成立しないことだった。一見単純な課題でつまずく一方、人間には相当に難しく見える課題を鮮やかに処理する。同じ知識労働のワークフローの中に、AIによって人間の能力が大幅に増幅される領域と、AIを使うことでかえって成績が落ちる領域とが隣接して存在する。これが jagged、すなわち鋸歯状という言葉の意味である。

Inside the frontier, the productivity gains were substantial. Participants using AI completed more tasks, worked faster and produced higher-quality output. Outside it, the direction of the effect reversed: when the task was deliberately designed so that AI would produce a plausible but incorrect analysis, users of AI became less likely to reach the right answer.

フロンティアの内側では効果は大きかった。AIを利用した参加者は12.2%多くのタスクを完了し、25.1%速く、成果物の質も改善した。一方、AIがもっともらしいが誤った分析へ誘導されるよう設計されたフロンティアの外側では、AIを使った参加者の正答率が低下した。AIは平均的に能力を上げる。しかし、境界の向こう側へ一歩出た瞬間、その補助が逆方向へ働くことがある。

The experiment used GPT-4 as it existed in 2023. It would therefore be a mistake to read its diagram as a map of AI capability in 2026. The frontier itself moves outward as models improve. The enduring value of the concept lies elsewhere: the boundary is not smooth, and the user never knows its geography perfectly.

もちろん、実験で使われたのは2023年当時のGPT-4であり、この図を2026年のモデル能力表として読むべきではない。フロンティアそのものはモデルの進歩とともに外へ動き続ける。だが、この概念の価値は「どの仕事までAIに任せられるか」という座標にあるのではない。能力境界は滑らかではなく、しかもその地形を利用者が完全には知り得ないという認識にある。

The coordinates have aged. The terrain has not.

古くなったのは座標であって、地形ではない。

what we expect — a smooth boundary looks hard — AI nails it looks easy — AI fails what is real — a jagged frontier (above the line = strong, below = weak) 期待する境界 — なめらかな線 難しそう — AI は得意 やさしそう — AI は苦手 実際の境界 — 鋸歯状のフロンティア(線の上=得意、下=苦手)
Fig. 2 — The Jagged Technological Frontier. The capability boundary of AI is not a continuous line. Tasks inside and outside the frontier can coexist within the same piece of work, while the frontier itself moves as models improve. What matters, therefore, is less the memorization of a static list of “tasks AI is good at” than the habit of asking whether one may already have crossed the frontier without noticing.図2 Jagged Technological Frontier。AIの能力境界は、一枚の滑らかな線ではない。同じ仕事の中にもフロンティアの内側と外側が混在し、その位置自体もモデルの進歩とともに動く。したがって「AIはこの仕事が得意だ」という固定表を覚えるよりも、自分がいまフロンティアのどちら側にいるのかを疑い続ける方が重要になる。

3As performance rises, assurance becomes more expensive性能が上がるほど、確かさは高くつく

This leads to the paradox at the heart of checking.

A low-performing system is, in some respects, cheap to verify. Its mistakes are frequent and obvious. A high-performing system creates a different problem: extracting the rare residual error while maintaining the same level of confidence becomes increasingly difficult.

ここから「確かめる費用」の逆説が生まれる。性能の低いシステムを検証することは、案外安い。誤りが頻繁で、しかも露骨だからである。一方、平均精度の高いシステムから稀な残存誤差だけを取り出し、同じ確信水準を維持しようとすると、検証は次第に難しくなる。

What rises is not necessarily the total amount of checking. As AI improves, the total human effort devoted to verification may well fall. What can rise is the marginal cost of assurance — the cost of obtaining the next increment of confidence that the output can safely be relied upon.

If ninety-nine outputs are known to be correct, manually reviewing the hundredth with the same intensity may erase the efficiency gain created by AI in the first place. Yet weakening the verification process means accepting the possibility that the hundredth output is the one hiding the tail event.

ここで上がるのは、単純な「確認作業」の総量とは限らない。AIの性能向上によって、確認の総工数そのものは大きく減るかもしれない。高くなるのはむしろ、一定の確かさを得るための限界費用である。99件が正しいと分かっているとき、百件目まで同じ強度で人間が読み直せば、AIによる効率化を自ら食い潰す。かといって検証を粗くすれば、百件目に潜むtailを取り逃がす。

Financial risk management has faced this trade-off for decades. If every portfolio were run every day as though the next Lehman failure were certain to occur tomorrow, the institution might become safe only by ceasing to be economically viable. Risk management therefore works through layers: ordinary measurement, limits, margin, capital, historical stress, hypothetical stress and escalation. Each layer addresses a different class of failure.

これは金融のリスク管理と同じ難問である。すべてのポジションを「次のリーマン・ショックが来る」と仮定して毎日運営すれば、システムは安全になる前に経済合理性を失う。だからリスク管理は、平時の測定、ヒストリカル・ストレス、ハイポセティカル・ストレス、限度枠、証拠金、資本といった異なる防護層を重ねる。バーゼルのストレステストの枠組みも、その結果だけでなく、前提、リスクの網羅性、モデルリスク、シナリオの妥当性そのものを理解する必要を強調している。

AI will probably require a similarly layered architecture of assurance. The answer is neither “let a human check everything” nor “the model is accurate enough, so trust it.” The relevant questions are where verification can be automated, where one must return to source data, where human judgment belongs, and which classes of failure are sufficiently consequential that a single occurrence should stop the system.

AIにも、おそらく同じ多層的な確かさの担保が必要になる。ただ「全部人間が確認する」でも、「精度が高いから信用する」でもない。どこを自動検証し、どこを一次データまで戻り、どこに人間の判断を置き、どの失敗には一件でも止めるべき重大性があるのか。その設計こそが検証の中心になる。

The scarce capability is moving from checking every answer to designing the checking regime itself.

4The more polished it becomes, the harder the error is to see綺麗になるほど、見えなくなる

The Jagged Frontier study contains another result that interests me even more than the headline numbers. Even when participants crossed the frontier and reached the wrong conclusion, the answers produced with AI were often rated more highly for coherence and persuasiveness.

Jagged Frontierの研究には、正答率以上に興味深い結果がある。フロンティアの外側で結論を誤った場合でさえ、AIを利用した参加者の回答は、文章の一貫性や説得力という点では高く評価された。

Key point要点 AI does not merely reduce the number of bad answers. It can make a bad answer look finished. AIは誤答をなくすだけではない。誤答まで、より完成された形で提示する。

That is a difficult property for human beings to manage. We instinctively distrust a clumsy argument. A document that is poorly structured, inconsistent or visibly incomplete invites scrutiny. A polished answer — logically ordered, numerically supported, apparently balanced, already anticipating objections — acquires a kind of borrowed authority.

As model quality improves, correctness and presentation quality rise together, and the visual distance between a correct answer and an incorrect one narrows.

これはかなり厄介である。人間は粗雑な文章を疑うが、論旨が整い、数字が入り、反論まで先回りした文章には権威を感じる。AIの性能向上は、正しさと同時に見せ方の品質も引き上げるため、正解と誤答の外観上の距離が小さくなっていく。

I notice a related phenomenon when writing. Pass the same paragraph through AI often enough and the redundancies disappear, the transitions improve and the reader is less likely to stumble. Yet over-editing can also remove the very irregularities that reveal where the writer is still uncertain. An unresolved tension can be transformed into an elegant paragraph before the underlying thought has actually been resolved.

The cost of producing something that looks true is collapsing.

私は文章を書くときにも、これと似た現象を感じる。AIに何度も推敲させれば、重複は消え、接続は滑らかになり、読者が迷う場所も少なくなる。だが、磨きすぎれば、最初の文章にあった異物感や、書き手自身がまだ答えを持っていない部分まで整形されてしまう。正しいかどうかとは別に、「正しそうに見えること」の生産コストが急速に下がっている。

For the person responsible for checking it, that is not necessarily good news.

これは、確かめる側には不利な変化である。

5“Human in the loop” is not enough「ヒューマン・イン・ザ・ループ」では足りない

One response is to place a human at the end of the process: AI produces the work, a human approves it. Human in the loop.

I do not think that is sufficient.

では、最後に人間を一人置けばよいのか。AIが作り、人間が承認する――いわゆるヒューマン・イン・ザ・ループである。私は、それだけでは弱いと思う。

The important question is not whether a human exists somewhere in the loop, but what that human has the expertise and authority to challenge. If someone simply reads the AI-generated conclusion and clicks approve because nothing looks obviously wrong, that person is not a control. They are the final click in the workflow.

問題は人間がループの中に存在するかではなく、その人間が何を疑う能力と権限を持っているかだからである。AIの結論を読み、違和感がなければ承認を押すだけなら、その人間は安全装置ではない。ワークフローの最後に置かれたクリックにすぎない。

Expertise matters when it can travel backward from the conclusion to the assumptions that produced it. In finance, the same headline “loss” can change materially depending on the observation period, liquidity horizon, haircut, correlation assumptions, netting treatment or close-out period. The specialist is not merely the person who calculates the number faster. The specialist understands the boundary conditions under which the number is meaningful.

専門知が効くのは、結論を読み直すときよりも、その結論を成立させている前提へ遡れるときである。同じ「損失」という数字でも、観測期間、流動性前提、ヘアカット、相関、ネッティング、クローズアウト期間をどう置くかによって結果は変わる。専門家とは数字を速く計算する人ではなく、その数字を成立させているモデルの境界を知っている人でもある。

That is close to the deeper lesson of the Jagged Frontier. Because the capability boundary of AI cannot be known perfectly from the outside, the response cannot simply be to draw a permanent line labelled “AI stops here.” The system must retain the ability to return to primary sources, separate the source of truth from the interpretive layer, introduce independent challenge, place deterministic controls around irreversible actions, and stop when behaviour moves outside expected bounds.

What matters is not the slogan human in the loop. It is verification built into the architecture itself.

これはJagged Frontierの図とよく似ている。AIの能力境界を外から完全に知ることはできない以上、重要なのは「ここから先はAI禁止」という固定線を引くことではない。一次資料へ戻れること、数値の正本を分離すること、別のモデルや人間に独立して疑わせること、不可逆な操作には決定論的な制御を置くこと、そして想定外の挙動が起きたときに止められること。必要なのはヒューマン・イン・ザ・ループという標語ではなく、検証可能性そのものをアーキテクチャへ埋め込むことである。

AI output fluent · plausible human questions the premise corrected judgment not the material AI runs on — the lens its output passes through AI の出力 流暢・もっともらしい 人間 前提を問い直す 補正後の 判断 AI を動かす素材ではなく、出力を通すレンズとして
Fig. 3 — Expertise as a lens through which output must pass. Expert knowledge should not be treated merely as additional material to be fed into AI. It is also the lens through which AI output is examined. And that lens must inspect more than the answer itself: it must look at the assumptions, provenance, exceptions and the consequences if the answer turns out to be wrong.図3 専門知を、出力が通過するレンズに置く。専門知は、AIへ投入する追加データではない。AIの出力を通過させるレンズでもある。そのレンズが見るべきなのは答えそのものだけではなく、その背後にある前提、出典、例外、そして誤りが顕在化したときの影響である。

6Do not undervalue the people who know how to verify確かめる人を、安く見ない

AI will continue to improve. The Jagged Frontier will move outward, and many tasks that require human checking today will eventually become machine-verifiable. I am not arguing that a permanent profession of human “checkers” will become increasingly valuable simply because AI exists.

AIはこれからも賢くなる。Jagged Frontierは外へ動き、今日人間が確認している仕事の相当部分が、明日には機械的に検証できるようになるだろう。だから「確かめる人という職業が永遠に高価になる」と言いたいわけではない。

The scarce capability is not checking in the clerical sense.

It is designing certainty.

価値が上がるのは、単純な確認作業ではない。確かさを設計する能力である。

Which number deserves a return to the primary source? Which level of error is tolerable? Which failure cannot be tolerated even once? Which known crisis should be replayed as historical stress, which future crisis should be constructed hypothetically, and what might still remain outside the model after both have been done?

Those questions require expertise, but also something harder to institutionalize: humility about the perimeter of one's own model.

どの数字を一次資料まで戻すのか。どの誤差は許容するのか。どの失敗は一件でも看過できないのか。ヒストリカル・ストレスで何を再現し、ハイポセティカル・ストレスで何を想像し、それでもモデルの外側に何が残っていると考えるのか。そこには対象についての専門知と、自分たちのモデルが何を見落とし得るかという謙抑が要る。

Generative AI has already made the search, organization, summarization and presentation of information dramatically cheaper. As that continues, the answer itself becomes less scarce. What becomes scarcer is a defensible reason for believing the answer.

生成AIによって、情報を探し、整理し、要約し、もっともらしい文章へ仕立てる費用は急速に下がった。これからは、答えそのものより、その答えを信じる根拠の方が希少になる。

Financial markets teach the same lesson in a harsher form. Risk cannot be driven to zero. At best, one can understand what risk is being taken, build defenses against what can be imagined, and remain conscious that something still lies outside the model.

AI will probably be no different.

金融市場では、リスクをゼロにはできない。できるのは、どのリスクを取っているかを知り、その外側にまだ何があるかを忘れないことだけである。AIも、おそらく同じだろう。

The Seed種としての一行 Answers will become cheap. Certainty will have a price. 答えは安くなる。確信には、値段がつく。

References参考文献

1.Fabrizio Dell'Acqua et al., “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality,” Organization Science, Vol. 37, No. 2, 2026, pp. 403–423. doi.org/10.1287/orsc.2025.21838Fabrizio Dell'Acqua et al., “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality,” Organization Science, Vol. 37, No. 2, 2026, pp. 403–423. doi.org/10.1287/orsc.2025.21838

2.Bank of England, Financial Stability Report, December 2022.Bank of England, Financial Stability Report, December 2022.

3.Federal Deposit Insurance Corporation, materials concerning the failure of Silicon Valley Bank and the 2023 regional banking turmoil.Federal Deposit Insurance Corporation, materials on the failure of Silicon Valley Bank and the 2023 regional banking turmoil.

4.Federal Reserve Board, Financial Stability Report, April 2025, and related analysis of the April 2025 Treasury-market turbulence.Federal Reserve Board, Financial Stability Report, April 2025; and subsequent analysis of the April 2025 Treasury-market turbulence.

5.Basel Committee on Banking Supervision, guidance on stress testing and liquidity-risk management.Basel Committee on Banking Supervision, Stress Testing and Principles for the Management and Supervision of Liquidity Risk.

The views expressed are the author's own and do not represent those of any organization he belongs to.掲載する内容は筆者個人の見解であり、筆者が所属する組織・団体の見解を示すものではありません。