﻿1
00:00:00,349 --> 00:00:03,450
AI reasoning ability is only 0.37% of a human's.

2
00:00:04,349 --> 00:00:04,849
Two weeks ago,

3
00:00:05,149 --> 00:00:07,549
a new test was quietly launched.

4
00:00:07,868 --> 00:00:09,990
It didn't make any major headlines,

5
00:00:10,108 --> 00:00:11,868
but within Silicon Valley's inner circles,

6
00:00:11,868 --> 00:00:13,269
it caused many sleepless nights.

7
00:00:13,669 --> 00:00:15,589
Because the results of this test show

8
00:00:15,669 --> 00:00:18,068
that today's most powerful AI,

9
00:00:18,068 --> 00:00:22,949
when faced with entirely new environments, has only 0.37% of human reasoning ability.

10
00:00:23,388 --> 00:00:24,730
Not 37%,

11
00:00:25,109 --> 00:00:26,550
but 0.37%.

12
00:00:26,908 --> 00:00:27,129
Hello,

13
00:00:27,349 --> 00:00:27,809
everyone.

14
00:00:28,068 --> 00:00:29,109
I'm Wang Lijie.

15
00:00:29,388 --> 00:00:32,009
Today I want to talk to you about a test called ARC-AGI-2,

16
00:00:32,109 --> 00:00:33,030
published on March 25, 2026.

17
00:00:33,308 --> 00:00:36,429
If you follow AI news,

18
00:00:36,668 --> 00:00:38,789
you know how crazy the pace of AI benchmarks has been in the last two years.

19
00:00:38,829 --> 00:00:42,329
One day a model tops a leaderboard;

20
00:00:42,908 --> 00:00:45,630
the next day, another model pushes it off.

21
00:00:45,668 --> 00:00:48,109
Every few weeks, someone announces AI has surpassed humans in another exam.

22
00:00:48,348 --> 00:00:52,549
Amidst this atmosphere of constant triumph,

23
00:00:52,829 --> 00:00:54,950
a test suddenly appears saying AI is only at 0.37% of human level.

24
00:00:54,988 --> 00:00:57,759
It sounds almost like a joke,

25
00:00:57,759 --> 00:00:59,189
but it isn't.

26
00:00:59,389 --> 00:01:01,189
It might be the most important "reality check" in the AI field in years.

27
00:01:01,389 --> 00:01:02,390
The creator of this test is François Chollet,

28
00:01:03,268 --> 00:01:07,209
a French AI researcher.

29
00:01:07,789 --> 00:01:10,870
He worked at Google for over nine years

30
00:01:11,108 --> 00:01:13,230
before leaving at the end of 2024 to start his own venture.

31
00:01:13,509 --> 00:01:15,829
He's no outsider;

32
00:01:15,948 --> 00:01:18,790
on the contrary,

33
00:01:18,948 --> 00:01:20,129
he is the author of Keras, the deep learning framework

34
00:01:20,549 --> 00:01:21,510
used by thousands of AI engineers worldwide

35
00:01:21,668 --> 00:01:24,390
to build neural networks.

36
00:01:24,588 --> 00:01:27,849
In other words,

37
00:01:27,989 --> 00:01:29,670
he's a builder of AI tools.

38
00:01:29,948 --> 00:01:30,569
Yet, it was this very "insider"

39
00:01:30,868 --> 00:01:32,590
who published a paper in 2019

40
00:01:32,868 --> 00:01:34,329
titled "On the Measure of Intelligence."

41
00:01:34,629 --> 00:01:36,790
The core argument of this paper can be summarized in one sentence:

42
00:01:36,909 --> 00:01:39,769
We have been measuring AI intelligence the wrong way.

43
00:01:40,388 --> 00:01:42,818
What does that mean?

44
00:01:42,828 --> 00:01:46,829
Think about it.

45
00:01:47,149 --> 00:01:48,290
All those AI report cards we've seen lately—

46
00:01:48,748 --> 00:01:49,670
exam scores, leaderboard rankings—

47
00:01:49,709 --> 00:01:52,250
what exactly are these tests measuring?

48
00:01:52,668 --> 00:01:54,308
They measure if AI can solve known types of problems,

49
00:01:54,308 --> 00:01:57,849
like doing math, writing code, or answering trivia.

50
00:01:58,349 --> 00:02:01,269
AI is indeed getting better at these,

51
00:02:01,709 --> 00:02:05,069
even surpassing humans in some cases.

52
00:02:05,509 --> 00:02:08,150
But Chollet says

53
00:02:08,389 --> 00:02:10,430
that this isn't intelligence at all—

54
00:02:10,669 --> 00:02:11,370
it's memorization.

55
00:02:11,588 --> 00:02:12,948
He makes a distinction

56
00:02:12,949 --> 00:02:13,949
using a classic psychological concept: fluid intelligence vs. crystallized intelligence.

57
00:02:14,508 --> 00:02:15,707
Crystallized intelligence is your accumulated knowledge and skills.

58
00:02:15,709 --> 00:02:20,229
Like speaking English,

59
00:02:20,549 --> 00:02:23,229
remembering the Pythagorean theorem,

60
00:02:23,269 --> 00:02:24,409
or knowing that Beijing is the capital of China.

61
00:02:24,628 --> 00:02:26,027
These are things you've learned, memorized, and can recall at any time.

62
00:02:26,028 --> 00:02:28,270
Fluid intelligence is completely different.

63
00:02:28,429 --> 00:02:31,949
It's your ability to spontaneously find a solution

64
00:02:32,669 --> 00:02:34,310
when faced with an unfamiliar, brand-new problem.

65
00:02:34,468 --> 00:02:38,909
As a simple example,

66
00:02:38,989 --> 00:02:42,110
imagine being dropped into a completely strange city.

67
00:02:42,549 --> 00:02:44,069
You don't speak the language,

68
00:02:44,348 --> 00:02:47,469
have no phone signal,

69
00:02:47,868 --> 00:02:48,530
and no map.

70
00:02:48,949 --> 00:02:49,810
Can you find your way back to your hotel?

71
00:02:50,308 --> 00:02:51,430
What would you do?

72
00:02:51,868 --> 00:02:53,569
You might observe local landmarks,

73
00:02:54,188 --> 00:02:54,889
pick a direction, and try it out.

74
00:02:55,788 --> 00:02:57,949
If it's wrong, you adjust based on new clues.

75
00:02:58,028 --> 00:02:59,990
Throughout this process, you aren't using pre-stored knowledge,

76
00:03:00,188 --> 00:03:03,150
but an ability for real-time perception, probing, feedback, and adjustment.

77
00:03:03,429 --> 00:03:07,389
That is fluid intelligence.

78
00:03:07,468 --> 00:03:11,909
Every normal person has it.

79
00:03:12,308 --> 00:03:13,789
[MISSING]

80
00:03:14,149 --> 00:03:15,250
[MISSING]

81
00:03:15,508 --> 00:03:16,949
And it's highly efficient.

82
00:03:17,308 --> 00:03:20,187
Chollet's core argument is that today’s AI,

83
00:03:20,188 --> 00:03:23,710
almost all of its amazing performance comes from crystallized intelligence.

84
00:03:23,949 --> 00:03:27,830
They have all the public text in human history as training data.

85
00:03:27,989 --> 00:03:31,229
Their "knowledge base" has long since surpassed any individual human.

86
00:03:31,429 --> 00:03:33,150
Ask it any known question,

87
00:03:33,308 --> 00:03:35,590
and it can give you a decent-looking answer.

88
00:03:35,788 --> 00:03:37,110
But that isn't intelligence.

89
00:03:37,149 --> 00:03:39,189
That's memory plus pattern matching.

90
00:03:39,669 --> 00:03:40,849
What is true intelligence?

91
00:03:41,269 --> 00:03:44,110
It's when you face an unprecedented situation,

92
00:03:44,269 --> 00:03:46,270
and the efficiency with which you learn and adapt.

93
00:03:46,389 --> 00:03:47,867
It's not about how much you know,

94
00:03:47,868 --> 00:03:49,590
but when you don't know,

95
00:03:49,748 --> 00:03:51,550
how fast you can figure it out.

96
00:03:51,868 --> 00:03:53,068
To prove this point,

97
00:03:53,068 --> 00:03:56,150
he designed the ARC-AGI test.

98
00:03:56,348 --> 00:03:59,310
The first version was called ARC-AGI-1,

99
00:03:59,468 --> 00:04:01,909
a series of abstract pattern transformation problems.

100
00:04:02,229 --> 00:04:04,750
You are given a few sets of input-output pattern examples,

101
00:04:04,829 --> 00:04:06,750
you must infer the underlying rules,

102
00:04:06,868 --> 00:04:09,789
and then apply those rules to a new input.

103
00:04:10,188 --> 00:04:10,530
For example,

104
00:04:10,788 --> 00:04:12,268
you see three sets of examples,

105
00:04:12,269 --> 00:04:15,389
each showing a blue square moving to a certain position.

106
00:04:15,508 --> 00:04:17,427
You need to find the movement rule

107
00:04:17,428 --> 00:04:20,029
and predict where the next square will move.

108
00:04:20,509 --> 00:04:21,509
The questions aren't hard;

109
00:04:21,629 --> 00:04:23,629
an average person can figure them out at a glance.

110
00:04:23,869 --> 00:04:27,067
But the key is that the rule for each problem is unique.

111
00:04:27,069 --> 00:04:28,709
You can't memorize answers in advance.

112
00:04:29,189 --> 00:04:31,329
It tests whether you can "get it" on the spot.

113
00:04:32,149 --> 00:04:33,189
The results were very interesting.

114
00:04:33,629 --> 00:04:36,750
Initially, AI performed poorly on this test,

115
00:04:36,949 --> 00:04:39,069
scoring only between 0% and 30%.

116
00:04:39,509 --> 00:04:42,990
Then AI companies began to focus their efforts on cracking this test.

117
00:04:43,269 --> 00:04:44,747
By late 2024,

118
00:04:44,749 --> 00:04:47,629
OpenAI released a reasoning model called o3.

119
00:04:47,749 --> 00:04:49,050
Using its highest compute mode,

120
00:04:49,468 --> 00:04:52,329
consuming 172 times more compute than the standard version,

121
00:04:52,869 --> 00:04:55,269
it achieved a score of 87.5%.

122
00:04:55,389 --> 00:04:58,589
Suddenly, it surpassed the 85% human average.

123
00:04:58,869 --> 00:04:59,949
Silicon Valley went wild.

124
00:05:00,309 --> 00:05:01,947
Many cheered,

125
00:05:01,949 --> 00:05:04,910
saying the dawn of AGI had arrived.

126
00:05:05,269 --> 00:05:06,939
But few noticed another number.

127
00:05:06,949 --> 00:05:10,709
The computational cost for this test was astronomical.

128
00:05:10,949 --> 00:05:13,949
Initial estimates were about $3,000 per problem,

129
00:05:13,988 --> 00:05:16,430
totaling over $1 million for 400 problems.

130
00:05:16,709 --> 00:05:18,788
Later, someone recalculated it

131
00:05:18,788 --> 00:05:22,148
and believed the real cost could be as high as $30,000 per problem,

132
00:05:22,149 --> 00:05:23,709
totaling over $10 million.

133
00:05:24,108 --> 00:05:25,250
Regardless of which figure is correct,

134
00:05:25,468 --> 00:05:26,750
it's terrifying.

135
00:05:26,949 --> 00:05:29,509
Meanwhile, an average human doing the same 400 problems

136
00:05:29,548 --> 00:05:32,550
uses energy roughly equivalent to eating one lunch.

137
00:05:32,749 --> 00:05:35,110
Human labor costs about $5 per problem.

138
00:05:35,389 --> 00:05:38,420
You spend millions or even tens of millions of dollars to tie

139
00:05:38,420 --> 00:05:40,029
with someone who just ate a sandwich.

140
00:05:40,069 --> 00:05:40,930
Is that a breakthrough?

141
00:05:41,588 --> 00:05:42,790
More importantly,

142
00:05:42,908 --> 00:05:46,269
when researchers analyzed o3’s process, they found

143
00:05:46,389 --> 00:05:48,430
it didn't actually "understand" the rules.

144
00:05:48,629 --> 00:05:50,588
It used brute-force search,

145
00:05:50,588 --> 00:05:54,290
probing through thousands of possible reasoning paths,

146
00:05:54,668 --> 00:05:56,230
backtracking, re-selecting,

147
00:05:56,309 --> 00:05:58,949
until it happened to find a path that worked.

148
00:05:59,309 --> 00:06:00,209
That's not thinking;

149
00:06:00,548 --> 00:06:01,670
that's exhaustive search.

150
00:06:02,028 --> 00:06:04,507
It's like trying to find the code to a combination lock,

151
00:06:04,509 --> 00:06:06,668
not by understanding the logic behind the code,

152
00:06:06,668 --> 00:06:09,550
but by trying from 0000 to 9999.

153
00:06:09,668 --> 00:06:11,069
One of them will eventually open it.

154
00:06:11,348 --> 00:06:12,829
You did open the lock,

155
00:06:12,869 --> 00:06:14,569
but do you really "understand" the lock?

156
00:06:15,348 --> 00:06:16,569
Seeing these results,

157
00:06:16,829 --> 00:06:18,069
Chollet wasn't discouraged;

158
00:06:18,108 --> 00:06:19,550
instead, he doubled down.

159
00:06:19,749 --> 00:06:22,149
He first launched ARC-AGI-2,

160
00:06:22,309 --> 00:06:24,069
significantly increasing the difficulty of the problems.

161
00:06:24,389 --> 00:06:28,550
AI scores plummeted from 93% to 68.8%.

162
00:06:29,108 --> 00:06:30,449
Then in March this year,

163
00:06:30,788 --> 00:06:32,670
he released ARC-AGI-3.

164
00:06:33,348 --> 00:06:35,870
This version completely changed the rules of the game.

165
00:06:36,228 --> 00:06:39,430
ARC-AGI-3 is no longer just static pattern puzzles;

166
00:06:39,509 --> 00:06:41,449
it features interactive environments.

167
00:06:41,949 --> 00:06:45,629
Think of them as miniature video game worlds.

168
00:06:45,908 --> 00:06:49,709
Each world is handcrafted by human game designers,

169
00:06:49,749 --> 00:06:52,550
with hundreds of completely different environments.

170
00:06:53,309 --> 00:06:55,750
But these games come with no manuals,

171
00:06:55,869 --> 00:06:57,230
no rule introductions,

172
00:06:57,309 --> 00:06:59,509
and they don't even tell you the objective.

173
00:07:00,028 --> 00:07:02,069
You are thrown into a strange world

174
00:07:02,149 --> 00:07:04,209
and must click and use trial and error,

175
00:07:04,588 --> 00:07:06,990
observing the changes caused by every action.

176
00:07:07,108 --> 00:07:10,850
Then you must figure out three things: how this world works,

177
00:07:11,228 --> 00:07:12,129
what the goal is,

178
00:07:12,389 --> 00:07:13,209
and what counts as winning.

179
00:07:13,829 --> 00:07:16,709
And it's not just about whether you pass;

180
00:07:16,829 --> 00:07:18,509
it also tracks how many steps you take.

181
00:07:18,749 --> 00:07:20,730
It calculates your learning efficiency—

182
00:07:21,108 --> 00:07:25,550
the speed at which you turn environmental info into effective strategies.

183
00:07:26,069 --> 00:07:30,018
This is Chollet's 2019 definition of intelligence:

184
00:07:30,028 --> 00:07:33,949
skill-acquisition efficiency over unknown tasks.

185
00:07:34,228 --> 00:07:36,709
This is the true test of fluid intelligence.

186
00:07:37,028 --> 00:07:38,990
It's not giving you a problem to solve;

187
00:07:39,269 --> 00:07:42,389
it's throwing you into a totally unknown environment

188
00:07:42,588 --> 00:07:46,149
to see if you can figure out what's happening from scratch.

189
00:07:46,548 --> 00:07:49,110
Humans scored 100% on this test.

190
00:07:49,509 --> 00:07:53,629
Every single human participant successfully solved every environment.

191
00:07:54,348 --> 00:07:55,930
As for the world's most powerful AIs?

192
00:07:56,548 --> 00:07:59,279
Google's Gemini 3.1 scored 0.

193
00:07:59,279 --> 00:07:59,990
37%.

194
00:08:00,389 --> 00:08:03,100
OpenAI's GPT-5.4 scored 0.

195
00:08:03,100 --> 00:08:03,870
26%.

196
00:08:04,189 --> 00:08:05,290
Anthropic's Claude

197
00:08:05,468 --> 00:08:07,209
Opus 4.6 scored

198
00:08:07,269 --> 00:08:08,670
0.25%.

199
00:08:08,869 --> 00:08:10,930
Most absurdly, Grok-4.20

200
00:08:11,269 --> 00:08:12,389
scored a flat zero.

201
00:08:12,629 --> 00:08:13,410
Listen carefully.

202
00:08:13,749 --> 00:08:15,470
It's not just that AI is slightly worse than humans;

203
00:08:15,548 --> 00:08:17,189
it's not a gap of 10 or 20 percent.

204
00:08:17,228 --> 00:08:19,389
It's a gap of two to three orders of magnitude.

205
00:08:19,709 --> 00:08:20,810
Humans: 100%.

206
00:08:21,189 --> 00:08:23,269
Strongest AI: 0.37%.

207
00:08:23,548 --> 00:08:25,509
If humans are standing at the mountain peak,

208
00:08:25,548 --> 00:08:28,509
AI hasn't even found the entrance at the foot of the mountain.

209
00:08:28,908 --> 00:08:29,810
Some might say,

210
00:08:30,269 --> 00:08:31,089
"That's not fair."

211
00:08:31,829 --> 00:08:35,230
These interactive environments are designed for humans.

212
00:08:35,349 --> 00:08:36,548
AI doesn't have hands

213
00:08:36,548 --> 00:08:37,289
or eyes.

214
00:08:37,629 --> 00:08:38,250
How can it operate?

215
00:08:38,828 --> 00:08:39,830
But the thing is,

216
00:08:39,869 --> 00:08:43,190
these environments run entirely on a computer screen,

217
00:08:43,269 --> 00:08:45,509
consisting of blocks and colors on a grid.

218
00:08:45,788 --> 00:08:46,889
The AI can see the screen and execute actions;

219
00:08:47,229 --> 00:08:48,429
it has the exact same input and interface as a human.

220
00:08:48,509 --> 00:08:52,789
It doesn't lack senses,

221
00:08:53,068 --> 00:08:54,610
computing power,

222
00:08:54,908 --> 00:08:55,868
or data.

223
00:08:55,869 --> 00:08:56,870
What it lacks

224
00:08:57,068 --> 00:08:58,710
is understanding.

225
00:08:59,229 --> 00:09:03,110
Even more intriguing is a detail from the preview phase:

226
00:09:03,149 --> 00:09:06,509
the top performer wasn't a trillion-parameter LLM.

227
00:09:06,589 --> 00:09:09,519
It was a relatively simple system built with

228
00:09:09,519 --> 00:09:11,230
reinforcement learning and graph search.

229
00:09:11,349 --> 00:09:13,470
It scored 12.58%,

230
00:09:13,589 --> 00:09:16,389
an order of magnitude higher than all the LLMs.

231
00:09:16,788 --> 00:09:17,529
What does this show?

232
00:09:17,908 --> 00:09:19,850
It shows that the LLM approach—

233
00:09:20,188 --> 00:09:23,889
learning statistics and patterns from massive text—

234
00:09:24,188 --> 00:09:26,148
is move-for-move the wrong methodology

235
00:09:26,149 --> 00:09:28,149
when facing truly novel problems.

236
00:09:28,349 --> 00:09:30,269
Language is not the sum total of intelligence.

237
00:09:30,749 --> 00:09:33,669
Feeding the entire internet's text into a system

238
00:09:33,749 --> 00:09:37,830
won't teach it to adapt through trial and error in a new environment.

239
00:09:38,229 --> 00:09:40,789
This reminds me of a very vivid metaphor.

240
00:09:41,068 --> 00:09:42,710
You know the story of AlphaGo.

241
00:09:42,948 --> 00:09:45,710
AI can defeat world champions at Go.

242
00:09:45,989 --> 00:09:47,788
Although Go has endless variations,

243
00:09:47,788 --> 00:09:52,110
the rules are fixed, known, and explained from day one.

244
00:09:52,389 --> 00:09:56,268
AI can achieve extreme optimization within this space of known rules,

245
00:09:56,269 --> 00:09:57,590
even surpassing humans.

246
00:09:58,269 --> 00:10:00,730
But what ARC-AGI-3 does is completely different.

247
00:10:01,188 --> 00:10:04,629
It's like being thrown into a board game you've never seen.

248
00:10:04,668 --> 00:10:06,227
No one tells you how the pieces move,

249
00:10:06,229 --> 00:10:07,750
or what counts as winning.

250
00:10:07,788 --> 00:10:09,830
You have to play a few rounds and figure it out yourself.

251
00:10:10,109 --> 00:10:11,590
Humans don't find this very difficult;

252
00:10:11,629 --> 00:10:13,470
we get the hang of it after a few tries.

253
00:10:13,629 --> 00:10:14,210
But AI?

254
00:10:14,708 --> 00:10:16,269
It becomes almost completely paralyzed.

255
00:10:16,428 --> 00:10:18,990
Consider a scenario from daily life.

256
00:10:19,308 --> 00:10:20,990
You visit a friend's house,

257
00:10:21,109 --> 00:10:24,668
and their toilet has a flush mechanism you've never seen—no button,

258
00:10:24,668 --> 00:10:25,330
no pull cord,

259
00:10:25,629 --> 00:10:27,789
but some strange rotating device.

260
00:10:27,989 --> 00:10:29,129
You might pause for three seconds,

261
00:10:29,389 --> 00:10:30,107
feel it out,

262
00:10:30,109 --> 00:10:30,570
give it a turn,

263
00:10:30,948 --> 00:10:32,029
and the water flushes.

264
00:10:32,188 --> 00:10:34,347
You never learned how to use that specific device,

265
00:10:34,349 --> 00:10:35,330
but you just "got it."

266
00:10:35,749 --> 00:10:37,990
You take this ability for granted.

267
00:10:38,068 --> 00:10:40,470
But for today's most powerful AI,

268
00:10:40,509 --> 00:10:42,230
this is the real challenge.

269
00:10:42,349 --> 00:10:45,467
It can write poetry, prove math theorems, and generate code,

270
00:10:45,469 --> 00:10:48,149
yet it can't handle an unfamiliar flush button.

271
00:10:48,389 --> 00:10:50,409
So, what do these experimental results really mean?

272
00:10:50,948 --> 00:10:54,429
If you watched my previous video on Penrose's theory of consciousness,

273
00:10:54,509 --> 00:10:57,149
you might already be connecting the dots.

274
00:10:57,509 --> 00:10:59,187
Starting from mathematics,

275
00:10:59,188 --> 00:11:01,830
Penrose used Gödel's Incompleteness Theorem to argue a startling conclusion:

276
00:11:01,840 --> 00:11:05,509
The human mind contains something "non-computable."

277
00:11:05,788 --> 00:11:07,508
There are things we can do

278
00:11:07,509 --> 00:11:10,230
that in principle, no algorithm can achieve.

279
00:11:10,469 --> 00:11:12,549
That was a purely theoretical deduction,

280
00:11:12,828 --> 00:11:15,860
whereas ARC-AGI-3 provides experimental

281
00:11:15,860 --> 00:11:16,750
evidence.

282
00:11:17,028 --> 00:11:17,970
It doesn't talk about consciousness,

283
00:11:18,149 --> 00:11:19,148
the soul,

284
00:11:19,149 --> 00:11:20,429
or quantum mechanics.

285
00:11:20,629 --> 00:11:25,309
It simply designs a simple test: in a brand new environment,

286
00:11:25,349 --> 00:11:27,570
can you learn to adapt from scratch?

287
00:11:28,149 --> 00:11:29,649
The result? All humans passed,

288
00:11:29,948 --> 00:11:31,509
while the AI failed across the board.

289
00:11:31,948 --> 00:11:33,049
A theoretical hammer,

290
00:11:33,269 --> 00:11:34,710
and an experimental hammer,

291
00:11:34,788 --> 00:11:36,707
hitting from two completely different directions,

292
00:11:36,708 --> 00:11:40,429
strike the same wall: there is something in human intelligence

293
00:11:40,548 --> 00:11:44,429
that cannot be replicated by stacking data, parameters, or computing power.

294
00:11:44,668 --> 00:11:45,769
This comparison is harsh,

295
00:11:46,068 --> 00:11:47,230
but it's also very real.

296
00:11:47,349 --> 00:11:51,268
We are so dazzled by AI's fluency in language and logic

297
00:11:51,269 --> 00:11:54,870
that we forget one fact: fluency does not equal understanding.

298
00:11:55,068 --> 00:11:58,347
A parrot can perfectly mimic every word you say,

299
00:11:58,349 --> 00:12:00,750
but it doesn't know what those words mean.

300
00:12:01,068 --> 00:12:03,750
Large language models are far more complex than parrots,

301
00:12:03,869 --> 00:12:06,649
but ARC-AGI-3 is asking the same kind of question.

302
00:12:06,659 --> 00:12:11,307
When you strip away all pre-trained knowledge and learned patterns,

303
00:12:11,308 --> 00:12:14,909
leaving only the ability to "figure it out on the fly when facing the unknown,"

304
00:12:15,028 --> 00:12:16,250
what does AI have left?

305
00:12:16,708 --> 00:12:18,870
The answer is: almost nothing.

306
00:12:19,229 --> 00:12:23,429
I want to emphasize: I'm not saying AI will never catch up.

307
00:12:23,668 --> 00:12:24,649
Technology is advancing,

308
00:12:24,948 --> 00:12:26,230
architectures are innovating,

309
00:12:26,349 --> 00:12:28,909
and perhaps someone will find a breakthrough next year.

310
00:12:29,109 --> 00:12:32,789
Chollet himself has set up a competition prize of $2 million

311
00:12:32,908 --> 00:12:36,000
for anyone who can reach human performance

312
00:12:36,000 --> 00:12:36,629
on ARC-AGI-3.

313
00:12:36,869 --> 00:12:40,190
But the gap revealed by ARC-AGI-3 today

314
00:12:40,229 --> 00:12:42,169
is not a simple quantitative difference.

315
00:12:42,548 --> 00:12:44,289
It's not that AI will catch up just by continuing to train

316
00:12:44,509 --> 00:12:46,750
or expanding its parameters.

317
00:12:47,469 --> 00:12:51,049
It exposes a qualitative difference: the entire current AI paradigm,

318
00:12:51,349 --> 00:12:53,690
from LLMs to reasoning-enhanced models,

319
00:12:54,028 --> 00:12:56,789
lacks a fundamental ability

320
00:12:56,908 --> 00:12:58,870
when facing true, absolute unknowns.

321
00:12:59,149 --> 00:13:00,649
What is this ability exactly?

322
00:13:00,989 --> 00:13:04,070
Chollet gives a technical answer: fluid intelligence.

323
00:13:04,188 --> 00:13:07,590
The ability to efficiently acquire skills in a completely new environment.

324
00:13:07,828 --> 00:13:09,669
But if you dig a step deeper,

325
00:13:09,749 --> 00:13:13,230
this answer just renames the problem.

326
00:13:13,509 --> 00:13:16,730
The real question is: why are humans born with this ability?

327
00:13:17,188 --> 00:13:20,909
If you put a three-year-old in a room they've never been in,

328
00:13:20,948 --> 00:13:21,730
without a teacher,

329
00:13:21,948 --> 00:13:23,029
or a manual,

330
00:13:23,188 --> 00:13:25,929
within minutes, they'll figure out how the doorknob works,

331
00:13:26,308 --> 00:13:28,830
which button turns on the lights, and where the blocks go.

332
00:13:29,028 --> 00:13:30,788
They don't need to train 100,000 times.

333
00:13:30,788 --> 00:13:33,549
They don't need to search through every possible combination of actions.

334
00:13:34,188 --> 00:13:35,049
They take one look,

335
00:13:35,269 --> 00:13:35,929
and just "get it."

336
00:13:36,548 --> 00:13:38,090
What exactly is this "getting it"?

337
00:13:38,629 --> 00:13:39,750
It's not computation.

338
00:13:40,028 --> 00:13:41,009
Because if it were,

339
00:13:41,269 --> 00:13:43,148
AI—being a trillion times faster than humans—

340
00:13:43,149 --> 00:13:44,429
shouldn't be worse at it.

341
00:13:44,788 --> 00:13:46,110
It's not a knowledge base,

342
00:13:46,188 --> 00:13:49,070
since a three-year-old has almost zero knowledge.

343
00:13:49,308 --> 00:13:51,548
Nor is it an instinctive response granted by evolution,

344
00:13:51,548 --> 00:13:54,710
since doorknobs and blocks weren't part of our evolutionary environment.

345
00:13:55,188 --> 00:13:57,169
It seems to be something more fundamental—

346
00:13:57,469 --> 00:14:00,250
an ability to instantly organize fragmented perceptions

347
00:14:00,269 --> 00:14:01,788
into a "meaningful whole."

348
00:14:01,788 --> 00:14:04,049
The ability to see "what this thing is for"

349
00:14:04,389 --> 00:14:06,070
without being taught.

350
00:14:06,668 --> 00:14:09,129
Cognitive scientists call this "Gestalt perception."

351
00:14:09,548 --> 00:14:12,710
But that's just labeling the mystery,

352
00:14:12,788 --> 00:14:14,669
not solving it.

353
00:14:14,948 --> 00:14:17,460
This is what's truly unsettling about

354
00:14:17,460 --> 00:14:18,149
ARC-AGI-3.

355
00:14:18,548 --> 00:14:20,750
It's not just saying AI isn't strong enough

356
00:14:20,828 --> 00:14:22,029
and needs further optimization.

357
00:14:22,308 --> 00:14:23,450
It asks a deeper question:

358
00:14:23,460 --> 00:14:26,409
What is "intelligence," anyway?

359
00:14:26,828 --> 00:14:30,227
We've always thought of intelligence as information processing: data in,

360
00:14:30,229 --> 00:14:31,067
process it,

361
00:14:31,068 --> 00:14:31,950
answer out.

362
00:14:32,188 --> 00:14:33,009
By that definition,

363
00:14:33,308 --> 00:14:35,429
AI should have outsmarted us long ago.

364
00:14:35,668 --> 00:14:37,429
But ARC-AGI-3 shows that,

365
00:14:37,469 --> 00:14:39,429
when facing the truly unknown,

366
00:14:39,509 --> 00:14:42,450
the human "biological computer," processing only 10 bits per second

367
00:14:42,788 --> 00:14:44,529
on 20 watts of power,

368
00:14:44,908 --> 00:14:48,590
demolishes silicon systems built with millions in electricity costs.

369
00:14:49,028 --> 00:14:49,370
Perhaps

370
00:14:49,668 --> 00:14:51,909
intelligence was never about information processing.

371
00:14:52,229 --> 00:14:53,668
Perhaps "understanding"

372
00:14:53,668 --> 00:14:56,470
cannot be attained by simply adding more computing power.

373
00:14:56,708 --> 00:15:00,429
Perhaps the moment a human understands something they've never seen,

374
00:15:00,469 --> 00:15:01,908
what occurs isn't a calculation,

375
00:15:01,908 --> 00:15:03,210
but a physical event—

376
00:15:03,509 --> 00:15:07,149
something current computer architectures simply cannot simulate.

377
00:15:07,828 --> 00:15:10,629
ARC-AGI-3 doesn't answer the question,

378
00:15:10,948 --> 00:15:14,307
but with that cold 0.37% figure,

379
00:15:14,308 --> 00:15:17,067
it pins the issue firmly to the table,

380
00:15:17,068 --> 00:15:19,870
making it impossible for anyone to ignore.

381
00:15:20,149 --> 00:15:22,730
Do you think AI can eventually bridge this gap?

382
00:15:23,028 --> 00:15:25,868
Or is the gap itself telling us

383
00:15:25,869 --> 00:15:28,570
that we've underestimated "understanding" from the start?

384
00:15:29,269 --> 00:15:30,409
If you have any thoughts,

385
00:15:30,548 --> 00:15:31,590
I'd love to hear them.

386
00:15:31,828 --> 00:15:32,490
I am Leo Wang.

387
00:15:32,668 --> 00:15:33,409
See you next time.
