Essay 010 · AI & Society

When AI Comes to Know the World AI が、世界を識る時

AI began with words — the easy part. The hard part was always the limbs. And the brain does not pilot all of them; the limbs fit the physics on their own. That is where what Japan still holds in its hand may yet matter. AI は言葉から始まった——易しい側だ。難しいのは、いつも手足だった。そして脳は、手足のすべてを操縦しない。手足が、自分で物理を擦り合わせる。日本がまだ手に持っているものが効くとすれば、そこだ。

June 2026 AI & Society AI と社会 9 min read 日本語 約9分

AI began with words — the LLM, the Large Language Model. From a vast past of text it predicts the next word, and writes, and summarizes, and translates. In very little time, that ability made its way into our ordinary days.

AIは、LLM、Large Language Model(大規模言語モデル)、つまり言葉から始まった。膨大な過去の文章から、次に来る言葉を予測し、書き、要約し、訳す。その能力は、わずかな時間で私たちの日常へ入り込んだ。

And the next AI is setting out to know the world: how things move, cause and effect in physics, gravity. From the state in front of it, it reads what comes next. On that ground, codified knowledge is not enough. What is needed is the knowing that accumulates through touching things, moving them, failing, and fitting them together. Japan has hoarded that in its hand, as craft, for a very long time.

そして次のAIは、世界を識ろうとしている。世界の動き、物理の因果関係、重力。いま目の前にある状態から、次に何が起きるのかを読む。その土俵で効くのは、形式知だけではない。物に触れ、動かし、失敗し、擦り合わせる中で蓄積されてきた知が要る。日本は、それを長いあいだ職人芸として手の中にためこんできた。

Say that, and the reply comes fast: Japan already lost the finished product. True. But step in one more pace, and another winning path appears. The key is the relationship between the brain and the limbs.

——と言うと、すぐ「日本はもう完成品で負けた」と返ってくる。その通りだ。けれど、もう一歩踏み込むと、別の勝ち筋が見えてくる。鍵は、脳と手足の関係にある。

1From Words to the World言葉から、世界へ

The LLM is a genius of words. About the world, too, it can say a surprising amount. But what a world model is reaching for is a little different.

LLMは、言葉の天才だ。世界についても、驚くほど多くを語ることができる。しかし、世界モデル(World Model)が扱おうとしているのは、それとは少し違う。

A cup tipping at the edge of a desk is going to fall. A car that starts to slide takes a different line. When a robot's finger touches an object, its hardness and its weight change the force applied next. From what is happening now, predict what happens next — because the world is not a pile of knowledge at rest but a thing that keeps moving into its next state. If the LLM reads the next word, the world model reads the next world.

机の端にあるコップが傾けば、このままでは落ちる。走っている車が滑れば、次の軌道が変わる。ロボットの指が物に触れれば、その硬さや重さによって、次に加える力も変わる。いま起きていることから、次に何が起きるかを予測する。世界は静止した知識の集積ではなく、刻々と次の状態へ移っていくからだ。LLMが次の言葉を読むなら、世界モデルは次の世界を読む。

Human beings do not know the physics of the world as language from the start either. And yet infants a few months old react differently than usual when an object suddenly disappears, or hangs in mid-air with nothing to support it.1 The expectation that the world ought to move a certain way is there remarkably early.

人間も、世界の物理を最初から言葉として知っているわけではない。それでも、生後数か月の乳児は、物が突然消えたり、支えもないのに宙に浮いたりすると、いつもとは違う反応を示すという1。世界には「こう動くはずだ」という期待が、ずいぶん早い時期から芽生えている。

By around the age of two,2 that turns from prediction by watching into a knowing that acts on the world. How to set a block so the stack will not come down. How to move a tool so the thing you want comes within reach. Touching, pushing, dropping, failing — the child takes cause and effect over to the body's side.

やがて二歳前後になると2、それは見るだけの予測から、自分で世界に働きかける知へ変わっていく。積み木をどう置けば崩れるか。道具をどう動かせば、欲しい物を引き寄せられるか。触り、押し、落とし、失敗する中で、因果関係を身体の側へ取り込んでいく。

I am not saying AI grows up the way a child does. Still, an intelligence that learned words next learns how the world moves, and then begins to learn how the world changes by its own action. There is something familiar in that order.

AIが子どもと同じように育っている、と言いたいわけではない。それでも、言葉を覚えた知能が、次に世界の動きを覚え、やがて自分の行為によって世界がどう変わるかを学び始める。その順序には、どこか既視感がある。

It begins to learn that the world has a continuation.

世界には続きがある、と覚え始める。

A world model, then, is less a matter of AI adding another layer of knowledge than of its coming to predict that continuation from the inside — or so I read it. Put that prediction into a body and it runs on into robots, self-driving, the arm on the factory floor. AI walks out from behind the screen.

世界モデルとは、AIが知識をもう一段増やす話というより、その「続き」を内側で予測できるようになる話なのだと思う。そして、その予測を身体に乗せれば、ロボット、自動運転、工場の腕へつながっていく。AIが、画面の外へ出てくる。

2The Brain Does Not Pilot the Limbs脳は、手足を操縦しない

When you run, the brain does not compute, step by step, "lift the right foot 3.5 centimeters, advance at 12 km/h." The brain does control movement, of course. But neither is it commanding every degree of freedom in the muscles, one by one, from the centre. Adapting to the tilt of the ground, the elasticity of muscle and tendon — part of the fine control is taken on by the spine, the periphery, and the physical properties of the body itself.3

走るとき、脳は一歩ごとに「右足を三・五センチ上げ、時速十二キロで」と計算してはいない。もちろん脳は運動を制御している。しかし、筋肉のすべての自由度を、一つずつ中央から命令しているわけでもない。地面の傾きへの適応、筋肉や腱の弾性——細かな制御の一部は、脊髄や末梢、そして身体そのものの物理特性が引き受けている3。

The body itself has become part of the computation.

身体そのものが、計算の一部になっている。

Robotics calls this computation by the body, morphological computation.4 The shape and the material of the body themselves take over part of the computation the brain would otherwise have carried. A soft finger conforms to the shape of what it holds; a springy leg absorbs the shock. A good body reduces what the brain has to think about.

ロボット工学では、これを身体による計算、morphological computation と呼ぶ4。身体の形と材質それ自体が、脳が担うはずだった計算の一部を肩代わりする。柔らかな指は物の形になじみ、ばねのある脚は衝撃を吸収する。良い身体は、脳が考えなければならないことを減らしてくれる。

And here Moravec's paradox bites. Machines handled checkers and intelligence tests — the things that look advanced to a human being — early on. Meanwhile "seeing," "grasping," and "walking," which a one-year-old does naturally, stayed hard for a long time. Behind the perception and the movement that look easy to us lies a thick knowing that a billion years of evolution carved into the sensory and motor regions of the brain.5

そして、ここでモラベックのパラドックスが効いてくる。機械は、人間には高度に見えるチェッカーや知能テストを早くからこなした。その一方で、一歳児なら自然にできる「見る」「掴む」「歩く」は、長いあいだ難題だった。人間には簡単に見える知覚と運動の裏側には、十億年の進化が脳の感覚野・運動野へ刻んできた、厚い知がある5。

So, apart from the words AI swallowed first, a thick world was still left over. The world on the side of seeing, touching, and moving.

つまり、AIが最初に飲み込んだ言葉とは別に、まだ厚い世界が残っていた。物を見て、触れて、動く側の世界である。

brain not every degree of freedom macro intent limbs — muscle · sensor · motor frictioncontractionelasticity — fit the micro-physics on their own — 脳 すべての自由度は命令しない マクロな意図 手足 — 筋肉 · センサー · モーター 摩擦収縮弾性 — ミクロな物理を、自分で擦り合わせる —
Fig. 1 — The brain does control movement, but part of the fine control is taken on by the spine, the periphery, and the physics of the body itself.図1 脳は運動を制御している。だが細かな制御の一部は、脊髄や末梢、そして身体そのものの物理特性が引き受けている。
reasoning, language — the thin skim (easy: AI drank it first) sense & motion — the thick base (hard: a billion years, in the brain's sensory and motor areas) — the defensible difficulty: the limbs — 推論・言葉 — 薄い上澄み(易しい:AI が先に飲んだ) 感覚と運動 — 厚い土台 (難しい:十億年、脳の感覚野・運動野に刻まれた) — 守れる難しさは、手足の側にある —
Fig. 2 — Moravec's paradox: reasoning is the thin skim; sense and motion are the thick, hard base.図2 モラベックのパラドックス。推論は薄い上澄み、感覚と運動が厚く難しい土台。

3So a Smart Limb Is Neededだから、賢い手足が要る

A similar division of labour is beginning to appear in physical AI. Up above, it thinks about where to go, what to grasp, what to do. Down where something actually walks and runs and grips the real world, it has to keep fitting itself to friction, to weight, to slippage, on a much faster clock. The brain does not need to compute every piece of the physics. It is better, in fact, if the limb can fit itself to the world to some degree on its own.

フィジカルAIにも、似た分業が現れ始めている。上では、どこへ行くか、何を掴むか、何をするかを考える。一方、現実を歩き、走り、物を掴むところでは、もっと速い時間で、摩擦や重さ、ずれに合わせ続けなければならない。脳が、すべての物理を逐一計算する必要はない。むしろ、手足の側がある程度、自分で世界に合わせられた方がいい。

A smart limb is not a limb that merely waits for fine-grained orders from the brain. If what it touches is hard it changes the force; if it slips it changes the grip. As it wears with repeated use, its own quirks change too. It senses that change and corrects the next movement a little.

賢い手足というのは、脳から細かな命令を待つだけの手足ではない。触れた物が硬ければ力を変え、滑れば握り方を変える。何度も動かすうちに摩耗すれば、自分自身の癖も変わっていく。その変化を感じ取り、次の動きを少し補正する。

However fine the brain, if the sensor measuring the force of contact is coarse, reality cannot be read accurately. If the motor does not move as intended, the intention does not reach the world. Heat, friction, wear, the quirks of a material — these are problems that did not exist inside the screen.

どれほど優秀な脳があっても、触れた力を測るセンサーが粗ければ、現実を正確には読めない。モーターが思った通りに動かなければ、意図は世界へ届かない。熱、摩擦、摩耗、素材の癖は、画面の中には存在しなかった問題である。

So the limb of physical AI is not the mere terminal of a command. It becomes the place where physics is fitted together, between the brain and the real world.

だから、フィジカルAIの手足は、単なる命令の末端ではない。脳と現実世界のあいだで、物理を擦り合わせる場所になる。

4This Is the Moatこれが、モートだ

Why might competitiveness remain here? Much of AI's code will, in time, be widely shared. But run machines in the real world and the same part, built to the same design, comes out a little different one unit at a time. Materials have quirks. There is heat. There is wear. There is manufacturing tolerance. And whoever has used that machine for a long time is left with a history of where it goes out of true.

なぜ、ここに競争力が残る可能性があるのか。AIのコードだけなら、いずれ広く共有されるものも多い。しかし現実世界で機械を動かすと、同じ設計の部品でも、一台ずつ少し違う。素材の癖がある。熱がある。摩耗がある。製造誤差がある。そして、その機械を長く使った者には、「どこでずれるか」という履歴が残る。

There are three walls. One is data asymmetry: the record of how a thing actually moved in the real world accumulates first where that thing was actually made and actually run. The second is the fit between hardware and software: load the same control onto a different body and it will not move the same way. Performance appears only when a precise body and a control that knows that body have been worked into each other. The third is time. A hand grows smarter by the number of times it has touched.

壁は、三つある。一つは、データの非対称である。物が現実世界でどう動いたかという記録は、その物を実際に作り、動かしたところにまず溜まる。二つ目は、ハードとソフトの擦り合わせだ。同じ制御を載せても、身体が違えば同じようには動かない。精密な身体と、それを知る制御が互いに合わせ込まれて初めて性能が出る。三つ目は、時間である。手は、触れた回数だけ賢くなる。

I have watched this from the side of financial infrastructure. Trades are matched, collateral is sized, delivery is made on the day it falls due. Most of that has long been the machine's work; no one checks it one item at a time. People remain, not to run the ordinary day faster. They remain because on the day something breaks outside what was anticipated, someone has to take the judgment. Carry the same rule over to another market and it will not work the same way — the rule and the quirks of the market it sits on have been fitted to each other over a long time. And only the side that has run the thing for years is left with a history of where it goes out of true: which route tends to clog, which number goes wrong first. It cannot be set down completely in a procedure manual, but whoever has run it knows.

私は、これを金融インフラの側で見てきた。取引を突き合わせ、担保の額を出し、期日どおりに受け渡す。その大半はとうに機械の仕事で、人が一件ずつ確かめてはいない。それでも人が残っているのは、平時を速く回すためではない。想定の外側で何かが壊れた日に、誰かが判断を引き受けなければならないからである。同じ規則を別の市場へ持っていっても、そのまま同じようには動かない。規則と、その規則が載る市場の癖とが、長い時間をかけて互いに合わせ込まれているからだ。そして、その仕組みを長く回してきた側にだけ、どこでずれるかという履歴が溜まる。どの経路が詰まりやすいか、どの数字が先に狂うか。手順書に書ききれるものではないが、動かしてきた者は知っている。

What has accumulated on the factory floor over many years comes down, pressed hard enough, to that as well. A slight difference in a material. A change with temperature. An odd noise. Wear. It cannot be set down completely on a drawing, but on the floor one knows that this one is about to go wrong. That accumulation is what we have long called craft. Craft is not skilled because it has hands. It is skilled because it has touched the world tens of thousands of times and remembered the answer.

長年の製造現場に蓄積してきたものも、突き詰めればそこにあるのではないか。素材のわずかな違い。温度による変化。異音。摩耗。図面には完全には書けないが、現場では「これはそろそろおかしい」と分かる。その蓄積を、私たちは長く職人芸と呼んできた。職人芸は、手を持っているから巧いのではない。何万回も世界に触れ、その返事を覚えてきたから巧い。

If so, what is genuinely hard to copy may not be the finished part itself. The part touches the world, goes slightly out of true, the drift is corrected, and it touches again. The source of the next intelligence is in that process.

だとすれば、本当に真似しにくいのは、完成した部品そのものではないのかもしれない。その部品が世界に触れ、少しずれ、そのずれを直し、また触れる。その過程にこそ、次の知能の源泉がある。

I wrote before that the land is free to till, and the land belongs to the lord. The view that whoever holds the base models and the semiconductors holds the land of AI has not changed. But go down into physical AI and the scenery shifts a little. The brain alone does not finish the work of the real world.

以前、「自由に耕せる、土地は領主のもの」と書いた。基盤モデルや半導体を握る者がAIの土地を持つ、という見方は今も変わらない。ただ、フィジカルAIへ下りていくと少し景色が変わる。脳だけでは、現実世界の仕事は完結しない。

You need not be king of the finished product. If the hand that touches the world becomes something the brain cannot easily replace, another kind of bargaining power appears there. If what Japan holds is going to matter, that is where.

完成品の王でなくてもいい。世界へ触れる手が、脳にとって容易に取り替えられない存在になれば、そこには別の交渉力が生まれる。日本が持っているものが効くとすれば、そこである。

the thinking limb — smart parts (Japan) base models · chips — the land (the lord) cannot be farmed without the hand 考える手足 — 賢い部品(日本) 基盤モデル · 半導体 — 土地(領主) 土地は、この手なしには耕せない
Fig. 3 — The land belongs to the lord. But if the hand that touches the world cannot easily be replaced, another kind of bargaining power appears there.図3 土地は領主のもの。だが世界へ触れる手が容易に取り替えられなくなれば、そこには別の交渉力が生まれる。

5But the Limb Must Be Made to Thinkただし、手足を、考えさせねば

Japan having precise limbs and those limbs having become "thinking limbs" are not the same thing. If what lies in the craftsman's hand disappears when the craftsman retires, it does not become an asset of the next era.

日本に精密な手足があることと、それが「考える手足」になっていることは、同じではない。職人の手の中にあるものが、職人の引退と一緒に消えるなら、それは次の時代の資産にはならない。

The tacit knowing in the hand has to be measured, kept, and moved into a form a machine can use. Run it, pick up the drift, return it to the next control. Today better than yesterday, tomorrow better than today, so that the hand itself comes to know the physics better. This is what I call the digitizing of the fit.

手の中にある暗黙知を、計測し、残し、機械が使える形へ移していく。稼働させ、ずれを拾い、次の制御へ返す。昨日より今日、今日より明日、その手自身が物理をよく識るようにする。これを「擦り合わせのデジタル化」と呼んでいる。

But that does not mean replacing the craftsman with AI. It means keeping, on the machine's side as well, a dialogue with the world that had been closed inside the craftsman alone. What the hand felt, and why it changed the movement. Only when that accumulation too can be handed to the next generation does the long accumulation of manufacturing carry into the age of physical AI.

ただし、それは職人をAIへ置き換えるという意味ではない。むしろ、職人の中だけに閉じていた世界との対話を、機械の側にも残していくことだ。手が何を感じ、なぜ動きを変えたのか。その蓄積まで次の世代へ渡せて初めて、長年の製造業の蓄積がフィジカルAIの時代へつながる。

So it is a little early to say that Japan already has a moat. The hand is there. A hand that has, for a long time, touched materials, run machines, and listened for the world's answer.

だから、日本がすでにモートを持っている、と言うのは少し早い。手はある。長いあいだ、素材に触れ、機械を動かし、世界の返事を聞いてきた手がある。

When AI moves from words to the world, what is needed is not only a large brain. It is limbs that see the present and read the next, that touch the world and remember its answer.

AIが言葉から世界へ移るとき、必要になるのは大きな脳だけではない。いまを見て次を読み、世界に触れ、その返事を覚える手足である。

The hand is still on this side.

その手は、まだこちらにある。

The Seed種としての一行 The moat is not ours yet. What we have is a hand that has spent a long time learning the world's reply. モートは、まだ持っていない。あるのは、長いあいだ世界の返事を聞いてきた手だ。

References参考文献

1.Renée Baillargeon et al., "Object Permanence in Five-Month-Old Infants," Cognition 20(3): 191–208, 1985. The "hangs in mid-air with nothing to support it" case: Amy Needham & Renée Baillargeon, "Intuitions About Support in 4.5-Month-Old Infants," Cognition 47(2): 121–148, 1993. Caveat: reading expectations about physics out of infant looking times remains contested (Linette Kunin et al., "Perceptual and Conceptual Novelty Independently Guide Infant Looking Behaviour: A Systematic Review and Meta-Analysis," Nature Human Behaviour 8: 2342–2356, 2024, supports the effect while reporting it is small).Renée Baillargeon et al.「Object Permanence in Five-Month-Old Infants」Cognition 20(3): 191–208, 1985。「支えもないのに宙に浮く」は Amy Needham & Renée Baillargeon「Intuitions About Support in 4.5-Month-Old Infants」Cognition 47(2): 121–148, 1993。留保: 乳児の注視時間から物理の期待を読み取る解釈は係争中(Linette Kunin et al.「Perceptual and Conceptual Novelty Independently Guide Infant Looking Behaviour: A Systematic Review and Meta-Analysis」Nature Human Behaviour 8: 2342–2356, 2024 は効果の実在を支持しつつ効果量は小さいとする)。

2.Rachel Keen, "The Development of Problem Solving in Young Children: A Critical Cognitive Skill," Annual Review of Psychology 62: 1–21, 2011. "Around the age of two" follows the range that paper groups under three years of age; it is a rough marker, not a transition point established in the literature.Rachel Keen「The Development of Problem Solving in Young Children: A Critical Cognitive Skill」Annual Review of Psychology 62: 1–21, 2011。「二歳前後」は同論文が3歳未満として一括りにする範囲に沿った目安であって、転換点として文献に定まっているものではない。

3.R. F. Ker et al., "The Spring in the Arch of the Human Foot," Nature 325: 147–149, 1987. Steve Collins et al., "Efficient Bipedal Robots Based on Passive-Dynamic Walkers," Science 307(5712): 1082–1085, 2005. Gerald E. Loeb, I. E. Brown & E. J. Cheng, "A Hierarchical Foundation for Models of Sensorimotor Control," Experimental Brain Research 126: 1–18, 1999.R. F. Ker et al.「The Spring in the Arch of the Human Foot」Nature 325: 147–149, 1987。Steve Collins et al.「Efficient Bipedal Robots Based on Passive-Dynamic Walkers」Science 307(5712): 1082–1085, 2005。Gerald E. Loeb, I. E. Brown & E. J. Cheng「A Hierarchical Foundation for Models of Sensorimotor Control」Experimental Brain Research 126: 1–18, 1999。

4.Rolf Pfeifer & Fumiya Iida, "Morphological Computation: Connecting Body, Brain, and Environment," Japanese Scientific Monthly 58(2): 48–54, 2005; Rolf Pfeifer & Josh Bongard, How the Body Shapes the Way We Think, MIT Press, 2006. Caveat: the definition and its quantification are disputed, and the "offloading" picture is explicitly rejected by Vincent C. Müller & Matej Hoffmann, "What Is Morphological Computation? On How the Body Contributes to Cognition and Control," Artificial Life 23(1): 1–24, 2017. It is used here as a grounded metaphor, not as settled neuroscience.Rolf Pfeifer & Fumiya Iida「Morphological Computation: Connecting Body, Brain, and Environment」Japanese Scientific Monthly 58(2): 48–54, 2005/Rolf Pfeifer & Josh Bongard『How the Body Shapes the Way We Think』MIT Press, 2006。留保: 定義・定量には議論があり、「脳の計算を身体が肩代わりする」像は Vincent C. Müller & Matej Hoffmann「What Is Morphological Computation? On How the Body Contributes to Cognition and Control」Artificial Life 23(1): 1–24, 2017 が否定している。本稿は確定した脳科学でなく根拠ある比喩として用いる。

5.Hans Moravec, Mind Children: The Future of Robot and Human Intelligence, Harvard University Press, 1988, p. 15. What was carved is the "large, highly evolved sensory and motor portions" of the brain, not the body; "one-year-old" and "intelligence tests or checkers" are from the same page. Caveat: the claim earlier in this section — that the body itself has become part of the computation — is not Moravec's but this essay's own, from a separate lineage in embodied robotics.Hans Moravec『Mind Children: The Future of Robot and Human Intelligence』Harvard University Press, 1988, p.15。刻まれた先は「脳の、大きく高度に進化した感覚野・運動野」であって身体ではない。「一歳児」「知能テストやチェッカー」も同ページ。留保: 本節前段の「身体そのものが計算の一部を担う」はモラベックの主張ではなく、身体性ロボティクスの別系譜に属する本稿の議論である。

The views expressed are the author's own and do not represent those of any organization he belongs to.掲載する内容は筆者個人の見解であり、筆者が所属する組織・団体の見解を示すものではありません。