Source. Transcript titled “Liang Wenfeng Investor Meeting · Audio Transcript.” Audio file
deepseek 0520.m4a, duration ~3 h 44 min. Recorded 20 May; compiled 2026-07-16. Automatic speech recognition plus AI cleanup; speakers not separated. Bracketed marks are audio timestamps. Proper nouns and figures may be wrong; the recording is authoritative.
Method. Extract only judgments Liang Wenfeng states, the grounds he gives, and the order in which he derives later claims from earlier ones. No imputed motives. No rewrite as an investment memo. Quoted sentences are the Chinese original with OCR spacing removed, then translated. Figures are transcribed as-is; ASR-risk items are listed at the end.
Bilingual. 中文版
1. Source and method
1.1 What this document is
An investor Q&A, not a refereed paper and not a speaker-approved written statement. The header states: automatic transcription, no speaker diarization, possible errors in names and numbers. Late in the session a host asks attendees not to leak sensitive figures (including accelerator counts) or share screen recordings. This brief still organizes the stated views for reading. Capacity, accelerator counts, payback periods, and profit multiples are transcript figures, not externally audited.
1.2 Extraction rules
| Rule | Practice |
|---|---|
| His judgments only | Investor questions, courtesy, and breaks are not claims |
| Claim vs ground | “What he holds” and “what he offers as reason” are kept apart |
| Hedge words kept | “I think,” “should,” “may” stay probabilistic |
| Terms not collapsed | “Continual learning,” “learn to learn,” and “continuous learning” are not forced into one technical definition |
| Figures not rounded | “Ten months,” “six times,” “20,000 H-equivalent” follow the transcript |
1.3 Hierarchy
He opens by saying later remarks will orbit vision, and that “this vision is real, not invented… otherwise you cannot explain many of the things we do.” This brief therefore treats “vision → restraint → AGI mainline” as the primary structure. Product, pricing, chips, organization, and financing are inferences on that chain, not a parallel strategy list.
2. The core causal chain
2.1 The sequence he reuses
The same chain is applied to open source, pricing, consumer products, enterprise sales, organization, and vertical integration.
flowchart TD A["Starting point: goodwill toward the world
not IPO or profit maximization"] --> B["Vision organizes people
vision = how you actually operate"] B --> C["AI is large enough
perhaps ~10% of human GDP"] C --> D["Monopoly is not feasible
exclusive capture would be discarded by history"] D --> E["Restraint is required:
open source / reasonable profit / no grab of adjacent businesses"] E --> F["Raises the probability of reaching AGI"] F --> G["AGI commercial value taken as given
share and monetization are secondary"] G --> H["Consumer / enterprise / API
are by-products on the mainline"] B --> I["Sole non-negotiable interest:
team stability"] I --> F
2.2 Each hop, with his stated ground
| Hop | Claim | Ground he gives |
|---|---|---|
| Origin | No founding intent to “make X money, go public” | The first few dozen people would not have joined if they thought that way; “very large goodwill toward the world”; “useful to humanity”; “something beyond money” |
| Vision | The most important thing in a company is vision; vision is not a wall slogan | Jack Welch: a large firm is managed by vision, not by rules; “vision is what you do, not what you say—how you actually run” |
| Organization | “We have no organization”; vision does the organizing | No KPIs; vision never written down; individual interpretations differ, the broad direction does not |
| Market size | AI may eventually take ~10% of human GDP | Contrast with software: a software market might be tens of billions of USD a year; open-sourcing it can shrink it to tens or hundreds of millions |
| No monopoly | “One party occupying this alone… you cannot occupy it alone” | “Objective law,” “view of history”: exclusive capture will be discarded; resistance and other blocking methods will appear |
| Restraint | “You need a mechanism that keeps the gains you can take limited” | “The more you think (GDP is ours), the less you can pull it off”; “the more you restrain, the more likely you can finish this” |
| Probability first | He does not doubt AGI’s commercial value; he prioritizes raising P(success), not taking more share | Spring Festival traffic: no chase of a super-app; “there are still watermelons later; what is in front may be sesame” |
| By-products | Consumer, enterprise, and API are not the goal | “On the path to AGI I have to pass this step”; “we are not doing this in order to do consumer or enterprise” |
| Core interest | Team stability is the largest, even the only, core interest | “If people do not leave… we will reach AGI”; money and resources “are certainly not the problem” |
2.3 His test of whether the vision is real
He repeatedly says that without this vision many actions cannot be explained: insisting on open source, internal cheering at price cuts, not turning Spring Festival users into a super-app, not doing video generation despite it being “good business.” He treats “it explains what we already did” as evidence that the vision is genuine, not as branding.
“This vision is real, not invented. We really think this way and really act this way. Otherwise you cannot explain many of the things we do.”
3. How vision organizes the firm
3.1 Claim
The founding motive is not profit maximization and not a capital-market path. The organizing device is not written rules or KPIs but an unwritten vision. He summarizes that vision as large goodwill toward the world and work that is useful to humanity and beyond money. It has never been written down. Interpretations may differ; the broad direction is shared.
3.2 Grounds
- Self-selection. The first few dozen people would not have come for money and listing.
- Management history. About twenty years ago he most admired Jack Welch. He now thinks most of what Welch said may be wrong, but “the most important thing in a company is its vision” still holds.
- Observable acts. Open source is intended (he contrasts Zhipu’s open source as feeling “forced”); people cheered in the company chat when prices were cut.
- Cohesion. AGI as a larger vision attracts stronger people, which he says yields a talent and organization advantage versus product-first firms.
“Vision is not a slogan on the wall. Vision is what you do, not what you say—how you actually operate.”
“This vision is not even written. It has never been written down as anything.”
3.3 Cost he acknowledges
“No organization” has benefits and costs; they will try to keep the former. Later he says headcount growth requires change: some groups need a stricter hierarchy, some stay loose and flat; “we have to make this adjustment immediately.”
4. Restraint as a strategy derived from a view of history
4.1 Claim
Restraint is not a moral ornament. He gives two layers: a historical claim—AI is too large, monopoly is theoretically inconsistent with the facts—and a strategic claim—forgoing some gains raises the probability of reaching AGI and buys ease (“we do not even need overtime”).
“You need a mechanism that ensures the gains you yourself can obtain are limited. Only then can you pull it off. Restraint is required.”
“For me, restraint is a strategy. Sometimes you give something up in order to get more of something else.”
4.2 “Those who take more will be beaten by those who take less”
A micro version of the same view:
- Even if a firm’s arithmetic of “5% of human GDP” theoretically works (he says OpenAI’s arithmetic can look workable), it can be beaten by someone willing to take 1%, then by someone willing to take 0.1%.
- “You need not actually take more. If your vision is to take more, you will be beaten by those whose vision is to take less.”
- Floor: take too little and the firm’s commercial logic fails. The landing point is reasonable return, not profit max and not zero profit.
- He treats China’s willingness to “take even less” as one source of challenge to US participants.
4.3 Restraint as a list of acts he names
| Act | What he says | Why he says they still capture value |
|---|---|---|
| Open source | Required by the vision; part of restraint | Section 5: no conflict with revenue at ~6× profit |
| API not priced to maximize profit | Ten-month equipment payback | Demand inelastic in that band; further cuts add little social value |
| Spring Festival users not turned into a super-app | No fight with ByteDance for users | “Watermelons later”; not pushing consumer last year “was probably right”—in hindsight the front was sesame |
| Adjacent businesses not seized | Does not want to be a rival to large or small internet firms; wants to enable them | “We have not received less of anything because of this” |
| Video generation / world models (at this stage) | Good business, little relation to the intelligence ceiling | “We do something only if it is on the intelligence roadmap” |
| No upstream vertical integration into apps | “I hope not”; “we only need to eat one piece” | AI is large enough; do the piece they are best at / see as core |
| Product lines, ads, e-commerce | Ads, e-commerce, local services discussed half a year ago “are certainly useless” | Change is too fast; product life cycles are short |
“I should not grab every sesame seed. … Looking back, not pushing consumer last year was probably right. You can see there really are larger watermelons later.”
4.4 The “ordinary people” narrative
He says the starting point was low: little money, few accelerators, no name, no drawing power—“just a group of very ordinary people.” He himself is “a university graduate, not from the most elite school.” He ties “ordinary people doing something uncommon” to restraint and vision, and uses it to explain how they accomplished some things “without weapons.” This is his own account, not an external citation.
5. Open source, pricing, and “six-times profit”
5.1 Open source: required by vision, also a commercial judgment
Claim. Open source is first a requirement of the vision; second, it helps make AI work commercially. He grants this “sounds contradictory” because historically open source conflicted with commercialization; he holds AI is different.
Grounds.
- Traditional software: a market of perhaps tens of billions of USD a year; after open source “perhaps only tens of millions or a few billion remain.”
- AI: may take ~10% of human GDP; cannot be occupied by one party.
- “It is not that if I do not open-source I can occupy the market. That is theoretically inconsistent with the facts.”
- Contrast: Zhipu also open-sources, but with “a sense of being forced”; “for us, this is the intended meaning.”
- “I cannot see the benefit of closed source”; “ByteDance’s model is closed. What is the benefit? I cannot see it.”
- Even with open weights the bar is high: others deploying and driving cost down is “very hard”; “it is not that once I open-source, they can easily match our deployment cost.” He calls this cost-control ability the “sweet point” of a firm of their size.
Whether the strongest model will be open. “We will open-source, and our strongest model may also be open-sourced.” The hedge is kept.
Whether open weights are weaker. “Is the open-source model the same as the one we deploy ourselves? It is the same.” They will not open a worse model and keep a better one. Ground: last year consumer was essentially open-sourced; “we did not see a conflict in the consumer service.”
Third-party deployment. He does not fear competition; he “only fears they fail to deploy,” with worse quality or higher cost. He wants to help the open-source community. Toward Alibaba, Zhipu, and Moonshot: willing to help them do better; “we do not lose anything.”
5.2 Pricing rule: ten-month equipment payback
Claim. API pricing corresponds to “reasonable profit,” not profit maximization. The rule: buy a batch of equipment on the market and recover cost in ten months. V3.2 Flash and others follow this. Financially, servers may be amortized over three or five years; commercially, ten months is enough.
He maps “ten-month payback” to “six-times profit”: “ten-month payback corresponds roughly to six-times profit.” Six times “looks high but is not”; it may later fall to 4× or 3×; further down “is probably not possible,” yet profit remains large.
“We only take a reasonable profit. It depends on your willingness, not on how large the profit is.”
5.3 Why not cut further, and why not raise to maximize profit
| Direction | Claim | Ground |
|---|---|---|
| Do not raise | Current price is not profit-max | “If we maximized profit we should set the price higher”; in this band “user demand has no elasticity”; doubling price barely changes token use, so revenue nearly doubles |
| Do not cut hard | Further cuts add little demand | “At this price everyone can afford it”; lower still “does not raise company revenue, and does not add much value for society” |
| Room remains | A comment said ten-month payback is “too profitable” | He agrees “there is still room to cut,” and room in model optimization; but “we can hit ten-month payback, others cannot,” and he says Alibaba or Tencent costs “should be several times higher”—his comparison, not their filings |
He tells an internal story: a model (transcript: DDCP, name uncertain) was priced high for fear of excess demand; the team was unhappy; he cut the price to one quarter; people cheered in the group chat. He treats this as evidence of the “real” intent: “very cheap, very good, so that everyone can use it fully.” He also notes that for competitors a price cut “is certainly not good news”; ARR falls.
5.4 When open source conflicts with revenue
No conflict if they take only ~6× profit. That price already leaves independent third-party deployment “unprofitable,” because third parties cannot match the cost.
Conflict if they wanted 100× profit: a third party deploying at perhaps 20× cost could still undercut 100×.
“The premise is that we only take six-times profit, recover cost in ten months… open source will not have much effect.”
He says this pattern is “sustainable in the long run” under their vision, or “we intend to do it this way.”
5.5 Consumer / enterprise / API as by-products
Claim. Consumer users, enterprise revenue, and the API are intermediate outputs on the AGI path. They will be done and done well, but most people in the company “do not see this as equal in importance to AGI.”
Grounds and examples.
- Consumer last year: “We did not put much effort into consumer; at one point we even did not want to maintain those users, but we could not drive them away.” The Spring Festival surge “was not in our script.”
- Keeping users is “pure cost,” but “may be useful later”; “if we can pick it up in passing, we will.”
- Enterprise ARR this year: if demand keeps expanding and more GPUs can be bought, “ARR of several hundred million USD is quite possible”; if AI can reach one billion USD, “cash flow could turn positive.” They “will do this,” but it is “not the first priority.”
- API: “On the path to AGI I must pass this step. Serving these techniques through the API is not extra work.” “A few people to maintain the API is enough; we do not even have customer service or sales… users come on their own.”
- Higher-plane advantage: doing lower-tier applications from an AGI technical height. Firms whose vision is serving consumer or enterprise users have advantages in product, service, and traffic, not in technology. “The favorable fact now is that model technique is the most important thing.”
Enterprise ceiling. Under “this generation of AGI / AI technique,” enterprise demand will grow fast but is finite; “in the end it is limited by demand, not by compute.” Ten-month payback means they would buy if there were uses; “there are not that many uses.” Demand grows as technique breaks through.
Closed loops. He observes that “those who really want a consumer closed loop are not necessarily better than we are; those who really want an enterprise closed loop may not be better either.” Then: “the more you want…” (the recording trails off; no extension here).
6. The AGI technical ladder
6.1 Long-term aim
“The company’s long-term vision, I think our target should be AGI.” Definitions of AI differ; that does not stop them treating AGI as the target.
6.2 Capability bound: already above humans given full context
Claim. Relative to this generation of technique, if a problem is clearly specified and the model is given complete context and instructions, “it already surpasses humans.” The premise is hard: people have decades of context; AI is better than people only inside limited context, and still cannot replace an employee.
Employee analogy. A hire takes about two months to learn the firm and the job; “call Xiao Wang over” is enough. Without those two months of learning, an AI must be told who Xiao Wang is, the role, location, how to find them, and what to watch. “You cannot give it all context, and it is not realistic.”
Next gap. Continual learning. The transcript also has “we are still short of learn-to-learn for the next step,” mixed in the same passage with “continual learning.” This brief does not force the phrases into one concept.
6.3 Ladder: each step uses the previous
flowchart LR LM["Language model"] --> CoT["CoT / chain of thought
last year"] CoT --> Ag["Agent
this year"] Ag --> CL["Continual learning
next bottleneck"] CL --> S["Self-iteration
he calls this a singularity
but gradual"] S --> E["Embodied intelligence"]
| Step | How he defines it | Why it is a step |
|---|---|---|
| Language model | Prior step | CoT uses it |
| CoT | Self-thinking raises the intelligence ceiling | Already “surpasses the top humans” on olympiad math and programming; at its ceiling it still does not reach AGI |
| Agent | Broader range, higher ceiling | Uses CoT; after the step is exhausted it still cannot replace an employee |
| Continual learning | Learn over a long period as a person does, without a very strong training dump each time | Same problem as “completing tasks”; after Agent, “the next bottleneck we can see” |
| “Singularity” | The model can do everything humans can, including developing its next version | He immediately qualifies: not a jump, a long gradual change; people are merely used to calling it a singularity |
| Embodied intelligence | Enter the physical world: housework, elder care | After self-iteration; “for a normal person, the need is not a computer” |
Why this order. Later technique can be built with earlier technique. “After the singularity, embodied intelligence need not be done by us.” Doing embodiment first is “bitter work.” “This roadmap is the easiest”; “we need not work overtime.”
First use of self-iteration. Without bodies, what he wants AGI to do is “help me iterate the next model.” With bodies, he wants it to iterate the next embodiment—the next robot.
6.4 Multimodality, search, scaling
- Multimodality. Important for consumer products; for the intelligence ceiling “it is a component, not the mainline itself,” “like search.” They will do it, and “V4 and later V4 versions will support native multimodality.”
- Scaling. He believes larger scale is better and unlocks more functions. What blocks scaling is compute, not lack of will. Model size is computed from resources; “it is not that this size is enough.” They have not hit a scaling wall. “When Silicon Valley says scaling has topped out, that is for Silicon Valley; for us in China we are still far from that.” Scaling includes data, model size, and training cost.
- Definition of the next generation. “It must have continual learning before it counts as the next-generation model.” Until then they can improve cost, quality, and speed.
7. Sole core interest: team stability
7.1 Claim
After listing many forms of restraint, he names the non-negotiable: “Our largest core interest is keeping the team stable… it can even be seen as the only core interest.”
“If I can keep the team stable, we will certainly succeed; we will certainly reach AGI. It is that simple.”
“Money is certainly not the problem. Resources are not the problem. Other factors are easy to obtain.”
Delay (half a year, a year) is acceptable; “it will not be that we cannot build it.” The largest risk is people leaving. After the recent financing, “the options people received are still quite large,” and the risk “has been substantially reduced.” Logic: if the most important and longest-tenured staff stay, others will not leave even with fewer options—“they did not come only for money.” Historical attrition is lower than peers, but “this is still our largest challenge, the only challenge.”
“Except for this, I think we can do without everything else; we can restrain.”
8. What he places off the mainline
8.1 AGI mainline only
External stance: “We only do the AGI mainline”—he lists GPT, CoT, Agent, and the like. The field is wide; what is not on the mainline is not done.
| Not done (this stage) | His judgment | Qualification |
|---|---|---|
| 3D, video generation | Little relation to the intelligence mainline | Commercially “good business”; after Sora everyone did it, small firms later cut it |
| World models | “At present also not closely related to the intelligence ceiling” | “I mean at the current stage”; the term is “not that clear”; many things can be called a world model |
| Ads / e-commerce / local-service insertion | Commercialization talk from six months ago “is certainly useless” | Change is too fast; talking product lines in the past three years “was a waste of time” |
| App-side vertical integration | He wants others to do it | “I do not want to eat everything” |
| In-house chips (default) | He hopes to buy chips at a reasonable price | Power-plant analogy: running a plant does not require building generators |
Most important now: “get AI training right.” That does not need world models, “not even modality”—narrowing the task range does not break the algorithm. Next: continual learning; then “it asks its own questions.” That roadmap contains neither world models nor video generation.
8.2 Timing of productization
“What is most worth doing now is still AGI… pushing the floor of intelligence up.” That has larger return than more product lines or more commercialization paths. His habitual question: “What has the largest return at this moment?” Product work is not that.
“We have always been commercializing; we have just not taken commercialization as the goal.” “The time to turn fully to commercialization should be far away.” “In any future we can now see, focusing on product at any moment is too early.” He wants commercial uses done by society and partners.
Investors: chosen carefully, aligned in interest, “least hostile to us, or most hoping we succeed”—“not everyone hopes we succeed, because we still harm many other people’s interests.”
8.3 Vertical-application priority
He has not yet thought through the most workable domestic commercial model; China and abroad need not match. Given current conditions, “the most reasonable move is to put full force into a general agent”; finance and medical agents have lower priority; “at this stage what matters most should still be a coding agent.”
9. Compute, the US gap, domestic chips
9.1 Gap versus the US: resources, not talent
Claim. “Our gap with the US is only one thing: resources.” On people there is “almost no gap”: the same pool somewhat randomly stays or goes abroad; “domestic talent is not scarce”; the base is large and replenished every year.
Nature of the talent gap. Less compute → fewer experiments → training of people is affected. “Every difference we see—talent, model capability, applications—can be treated as a difference in compute resources.”
Two sources: domestic buyers cannot get that many accelerators; capital expenditure is lower than in the US. Headcount pay is a small share; “the bulk is still compute.” US “salaries of a hundred million USD of that kind”—transcribed as stated.
9.2 Compute and scale figures in the transcript
All figures below are ASR text, not verified.
| Item | Transcript figure | How he uses it |
|---|---|---|
| Current compute | ~20,000 H-equivalent; most arrived in the last one or two months; more machines may still be inbound | Last year was smaller; this year they are “expanding very aggressively”; coming months, large batches, “basically NVIDIA” |
| Buying rule | At a reasonable price, buy as many as can be bought; same after the financing is spent | “If I spend all the money within six months, I think that is a good thing”; cash earns about “two percent”; cards correspond to “ten-month social cost” |
| Premium | Willing to pay some premium to turn cash into cards | “Too good a deal”; still hard to buy that many |
| Procurement scale | “If I can spend 20 billion this year,” procurement has done “extremely well” | Aspiration, not a commitment |
| Largest models | ~800B active; China still at tens of B active, an order of magnitude off | “Even spending 50 billion we still cannot train them” |
| To train a model that large | 50,000 GB300, or 200,000 Huawei 950; “training only, research not included” | To state the gap, not a purchase plan |
| Near-term experiments | Tens of B active | Still “quite far” from training 800B |
| Next scale | With more resources, 150B, 156B, or 250B active | Competing with the US at that scale is “not even considered” now |
| Narrative target | Current story: behind the US by about 6–12 or 12–18 months (he also says “simply, two years”), using about 1/20 the compute | He wants to rewrite it: a few-fold less compute, time cut to 6 months or 3 months; selective lead is possible; overall lead is “not realistic” |
Strategy: first do well at tens of B active, a scale they can train and serve. A larger model “can be trained,” but they cannot research enough before training it.
“Being behind means you have more time” and can use more ingenious methods—the recording is cut; the later narrative is “one to two years behind, 1/20 the compute.”
9.3 Domestic chips, CUDA, TileLang
Claim. Domestic accelerators were hard to use because the software ecosystem was weak; NVIDIA CUDA’s moat “is being dismantled quickly.” Three reasons:
- With AI, “we can use AI to build the ecosystem.”
- New technique, named TileLang: writing CUDA kernels in that high-level language “can rewrite NVIDIA’s whole ecosystem quickly”; with AI, “there looks to be little obstacle,” but “it is not finished.”
- CUDA grew from gaming GPUs; the compute-card market is now larger than the gaming-card market, so “there is no reason the two still need to be coupled.” Dedicated chips (Huawei or NVIDIA itself) will unbind from CUDA.
Fact to verify within a year. “The ecosystem of domestic chips is entirely fine.” After that the only issue is capacity. If NVIDIA cards can be bought, substitution is hard; if not, “everyone is forced.”
Huawei. Main partner. Transcript: about 16,000 cards of capacity allocated to them; internet majors perhaps over 100,000. 16,000 950s “equal only 4,000 B-series,” “not enough to train a next-generation model, only the current generation.” The point of buying 950s “is still to help Huawei get this ecosystem right.” “Internet majors need it more. For us, we can buy some non-compliant cards.” Quoted as transcribed; no legal inference.
Substitution (transcript). Huawei 950 supernodes “can fully substitute” GB200 and GB300 in performance and price; more expensive but “the premium is limited”; “100% more expensive can already be treated as price substitution.” Cost: four Huawei cards for one NVIDIA card, and two years behind. 950 supernodes are Q3 or Q4 this year; GB200 was Q3 two years ago. Conclusion: “there will no longer be an ecosystem gap, but on the chip it is 4× plus two years.”
V3 training. NVIDIA hardware, “already not NVIDIA’s ecosystem”: they first wrote TileLang, “almost no dependence on NVIDIA’s ecosystem.” Repeating that process on Huawei cards completes the move.
TileLang and inference efficiency. Asked whether inference efficiency falls sharply: “It raises efficiency, substantially.” Hardware execution loss “of 1% to 2% is acceptable.” A separate project “uses AI to write TileLang”; all TileLang is still human-written, already much faster than writing CUDA.
Clusters / in-house chips. Clusters “are certain”; “all our clusters are built by us.” In-house chips depend on return; he hopes not to need them if chips can be bought at a reasonable price.
9.4 Depreciation (answer on “stale compute in three years”)
“NVIDIA cards you can basically depreciate over five years. Huawei cards at most three.” 950 “is still fine this year, next year still OK; after that it may really draw too much power.” Shorter life because “it is already two years behind NVIDIA,” but “the gap is not that large.” B200: “however many you can buy now, I think it pays.”
9.5 China’s role in the global division of labor
“Chinese firms are likely to play the role of largest output.” Largest capacity, most electricity; “our AI may end up as one of the three bodies.” China will make the product cheapest; in quality “many Chinese goods already do not differ much from US ones”; cheapness “may be systematically lower.”
Structural advantages versus the US: he names only cost and product / user experience. Cost: the other side “does not need to do it, so they do not develop the capability.” Product: domestically “many native firms, product capability is still fine.” Other dimensions: “if there is a structural advantage, I think perhaps there is not.”
10. End-state competition
10.1 Three end-state differences
“Where will the final gap among large models show? I think in the end there may not be a large gap.” If there is, three sides: cost, time, user experience. Cost first (BYD battery analogy: matching quality at that price is not easy); time second (a few months early or late differs); experience has some stickiness but “may not be essential.”
Model quality “must be compared at the same cost.” Good versus poor “should not sit in one specific stage; it should be overall.”
Industry implication: “no one should have windfall profit”; “those who control cost well earn a bit more, those who control it poorly a bit less. That is all.” Too many Chinese base-model firms; the US “perhaps has three.” Scattered resources are waste; convergence to “three or four competitors is already ample,” “prices will already support a price war.” Very high margins “cannot fit objective law.” Talent shortage is cyclical, resolved in about two or three years by training; “history has never had a long-run shortage of one kind of person.”
10.2 Anthropic / OpenAI / Google
Anthropic overtaking OpenAI “I think is not long-run; it is certainly extreme.” “OpenAI and Google will probably rise in turns.” Anthropic’s Code Agent lead “is not that large,” “not a crush of OpenAI.” Internally “perhaps half the people… think OpenAI is better.” Anthropic has a first-mover edge that “should be gone soon.” All three are strong; the most efficient “should have burned the least money”—he does not name which.
10.3 Why they care about computational efficiency
He observes that a commercial firm “has little incentive to chase model efficiency”: low cost means “what do you still charge.” Startups: “you do not hear them say low cost is what they pursue.” For them: part of the vision; “our colleagues are ordinary people” and have empathy; under scarce Chinese compute they want something affordable that can run on domestic cards. Lower cost also lets them train a larger model on limited compute. Large firms “can add resources”; they “put cost efficiency first.”
11. Organization and research culture
11.1 Decisions: seek consensus
“The company as a whole is built on consensus. It is not that I decide everything.” Authority and influence rest on consensus. Guidance “is very limited”; “only if there is consensus can I push, and only then will I push.”
11.2 Two tracks: formal work not above half
- Top-down (“formal”). Collective tasks, e.g. shipping V4, requiring company-wide coordination.
- Bottom-up. Each person does what they want; “nobody manages them, no KPIs.”
- Standard: “formal work had better not exceed half”—half the time is unassigned, for exploration by one’s own judgment. If compute can support it, or almost no compute is needed, no coordination is required.
11.3 Two reasons they rarely work overtime
- Research needs a relatively loose setting; tight pressure blocks research.
- “We are very focused… there is little we have to do.” This follows restraint. “You can see many of our products are incomplete; we have not gone to fill the gaps.” He calls this a culture.
11.4 Not Bell Labs; still a firm
No imitation target. “Each step is from actual conditions, seeking truth from facts.” Unlike the US labs, which “at least explore from the angle of a commercial firm.” Unlike Bell Labs: Bell Labs “explicitly did not need commercialization”; “we explicitly must commercialize… the government will not give me a cent.” “Enterprise will certainly matter to us, because we may later have to live on it. It does not matter now, because it is now a cost line.” “We are still a company in essence,” with choices about which money, when, how much, and by what means. Historically, firms with aims beyond profit found that those aims “did not hurt commercialization; they made it better.”
After headcount growth: some units get a stricter hierarchy, some stay flat; “we are already making this adjustment.”
12. Continual learning, data, hallucination
12.1 Continual learning: a problem, not one technique
Investors see Agent most; “for us as researchers, what we see more is learning.” Continual learning “may not be one technique; it is a problem,” with many techniques. “The whole world has not yet found a method that really works… still exploration.” Internally there are “ideas that look promising, but none has been made to work.”
Relation to recursive improvement: a questioner pairs continual learning with the foreign topic recursive improvement. Liang answers “no method that really works yet,” and does not define the two as the same term.
Why it must come first. Agent capability is limited because it cannot learn continually in an effective way; after that, AI capability would be very strong and would raise their own research efficiency a great deal. “If continual learning is done first, general intelligence may be easy”; otherwise “doing general intelligence by hand” is data- and labor-intensive, “poor cost-performance.”
First user is themselves. “The first goal of the models we build is not that others use them well, but that we use them well.” Use the model to raise DeepSeek’s speed on the next version. “We do need AI to help us reach AGI. It does not yet work autonomously; it still pairs with people. But it is already very useful.”
Resource allocation (“lottery-style exploration”). How to split people between a relatively definite scaling goal (the question contains Office / MIS / MILES; ASR unclear) and unsolved continual learning. Answer: this research “does not consume accelerators, only very few; it needs ideas”; “it is not a project; many people need to think about the problem.” “We call this ‘lottery scratching.’ The bar is low; anyone can try.” What differs from other firms is that they “spend time discussing this… treat it as important,” but “need not spend many resources.” Training models and efficiency experiments consume people and accelerators.
Taste and intuition. Asked whether taste still matters after self-evolution. “AI does not lack taste and intuition now. What it lacks is continual learning.” “Ask it to write an essay: I do not think its taste and intuition are a problem.”
12.2 Data
“Data should be almost half the model.” High-quality annotation is expensive; “US annotation cost and Chinese annotation cost have little difference”; on high-end data “there is no cost advantage.” Their capital structure “cannot support that much high-quality annotation.” Two tracks: annotate the cheap part first. “About half the people in the company are annotating. Half of the core researchers, the most important people, half are annotating.” “Solving AI at this stage depends on annotating data.”
After financing they will raise post-training spend, but high-quality annotation “is typically not a capital input”; the bottleneck is time. Domestically “one can treat this as having started only in the last half year”; “even without more capital, existing capital is enough to expand at the fastest speed,” but speed has a cap, “not money, not accelerators.” Within a year, high-quality data “done reasonably well… should be expectable domestically.”
Can models exceed human knowledge already spoken. He thinks yes. Ground: AlphaGo “played a move humans had never seen.” There may also be a ceiling; “we cannot see that limit now.” “Roughly, it can exceed on the basis of knowledge humans already have and can already state.” Real versus simulated data: “There cannot be only one method; there are many.”
12.3 Hallucination
“Hallucination can be treated as solvable, improvable, through better post-training.” People have not put much effort in. “For us hallucination is a problem, but we class it as perhaps a product problem. We will fix it, but it is not the priority.”
13. Timelines, release cadence, capital path
13.1 How long to AGI; can domestic hardware catch up
- AGI: no critical point, but nonlinear. He believes “AI can accelerate AI research,” so later progress may be nonlinear. At minimum “it has to be able to learn continually.”
- Catching foreign models under “the current paradigm”: “domestically one or two years should be enough to be close, or perhaps this year we can substitute for foreign models.” “It is still not AGI.”
- Domestic hardware: several years. First the ecosystem (“ecosystem is a confidence problem”), then capacity. He does not believe they will still be stuck on capacity in five years; “this year, next year, the year after… may still be stuck on capacity.”
Language-model scaling: “I do not currently see an upper bound.” The US can train 800B active and “cannot really use it… it is indeed a bit expensive”—his judgment.
13.2 Release cadence and 50B / 150B
“A comfortable release cadence for me is about one version every two or three months. The last release was perhaps late April, so the next may be late June.” “Each version should still be better than the last.”
At 50B active, versus “this wave of open source,” reasoning speed and quality “will not differ much.” Versus “a model they have not released,” the gap is larger and needs a bigger size, “for example 150B.” On then-current training progress, “optimistically we can start training by year-end; at least next April…”—the sentence is cut. The comparison target is transcribed OCE. A question mapping 150–250B to “O4.7” is not completed in the visible answer (he turns to TileLang). An online version is transcribed GCV4, “still quite rough.”
13.3 Research and the capital market
“I think we should be able to do both.” If this year they have “several hundred million USD of enterprise revenue” plus consumer users, there is already some commercial base; if next year’s enterprise demand grows again, “the company is not far from net profit.” “In the worst case, selling the API may be enough to support a listed company”—if technique freezes, they sell API with full force and do the service well. “We want a larger aim, but we also have a floor of results we can show.” Reason: “a high-leverage place, and a very fast field; it may just not be that hard.”
14. Claim table
Rows follow his causal order. The ground column uses only reasons he gave in the meeting.
| # | Claim | Ground given in the meeting |
|---|---|---|
| 1 | Founding intent is not money, listing, or profit max | Early self-selection; goodwill and “beyond money”; vision explains open source and cheers at price cuts |
| 2 | Vision organizes people, not KPIs or written slogans | Welch; vision = actual operation; never written |
| 3 | AI is too large to monopolize; exclusive capture is discarded by history | Perhaps ~10% of human GDP; contrast with software markets |
| 4 | A mechanism must limit own take; restraint is strategy | “The more you try to monopolize, the less you succeed”; restraint raises P(AGI) |
| 5 | A vision of taking more loses to a vision of taking less | 5% vs 1% vs 0.1%; floor is that the firm must live |
| 6 | Open source is required by vision and is a commercial judgment | No necessary benefit to closed source; deployment-cost bar; no consumer conflict last year |
| 7 | Reasonable profit ≈ ten-month equipment payback ≈ 6× | Inelastic demand in that band; own cost others cannot match |
| 8 | At 6×, open source does not hit revenue; at 100× it does | Third parties at ~20× cost cannot enter 6×, can enter 100× |
| 9 | Consumer / enterprise / API are by-products on the AGI path | Users with little spend; API with almost no sales or support; higher-plane advantage |
| 10 | Not chasing a super-app at Spring Festival is restraint | Sesame vs watermelon; in hindsight, not pushing consumer last year was probably right |
| 11 | Sole core interest is team stability | If people stay, AGI is a matter of time; money and cards are not the hard constraint |
| 12 | Ladder: LM → CoT → Agent → continual learning → gradual “singularity” → embodiment | Each step uses the last; later steps can be sped by earlier ones; embodiment first is bitter |
| 13 | No video generation or world models at this stage | Little relation to the intelligence ceiling; video generation is good business, not mainline |
| 14 | Multimodality is a component, not the mainline | Analogy to search; native multimodality still from V4 |
| 15 | Scaling is believed; the wall is far; compute is the bind | Size is computed from resources; Silicon Valley’s “topped out” is not China’s |
| 16 | US gap is mainly resources / compute | Talent is randomly split; experiments follow card count; 800B vs tens of B |
| 17 | Buy cards at a reasonable price until money is gone; spending out in six months is good | Cash ~2% vs cards’ ten-month social cost |
| 18 | CUDA moat is falling; TileLang + AI codegen | Gaming and compute uncouple; V3 already almost independent of the NVIDIA ecosystem |
| 19 | Domestic ecosystem can be shown within a year; capacity remains | No NVIDIA → forced domestic; Huawei 950: 4-for-1, two years behind |
| 20 | End state compares cost, time, experience | Battery-price analogy; experience not essential |
| 21 | Anthropic’s Code Agent lead is not durable | About half internally prefer OpenAI; the three alternate |
| 22 | China may be “one of the three bodies”: capacity, power, cheaper | Analogy to systematically cheaper Chinese goods |
| 23 | Product lines, ads, e-commerce are too early | Three years of experience; short life cycles |
| 24 | Vertical apps: coding agent first | Domestic commercial model not yet clear; finance and medical lower priority |
| 25 | Formal work not above half; little overtime | Research needs slack; restraint means fewer tasks |
| 26 | Not Bell Labs; must live commercially | No government money; enterprise may be the later livelihood, now a cost line |
| 27 | Continual learning not yet solved; lottery-style exploration; first user is themselves | No global method that works; ideas do not burn cards; AI to speed the next version |
| 28 | Hallucination can be improved by better post-training | A product issue, not the research priority |
| 29 | Data ≈ half the model; annotation bottleneck is time, not capital | US and China annotation costs similar; about half of core researchers annotate |
| 30 | Base-model firms will converge to about 3–4 | High margin does not fit the law; US has about three |
| 31 | Talent shortage is cyclical (~2–3 years) | Websites, pilots as historical analogies |
| 32 | This year or 1–2 years: substitute foreign models under the current paradigm, not AGI | Current method “is not very hard”; AGI at least needs continual learning |
| 33 | Comfortable cadence: one release per 2–3 months | Last ~late April, next ~late June |
| 34 | 50B close to current open models; unreleased large models need ~150B | Optimistically start 150B training by year-end (sentence cut) |
| 35 | Research and capital markets can both be required | Several hundred million USD enterprise + consumer users; worst case, API sale can support a listing |
| 36 | TileLang raises efficiency a lot; ~1–2% hardware-execution loss is acceptable | High-level language, small code volume, can be rewritten; also a project of AI writing TileLang |
15. ASR-risk items and term map
15.1 Symbols to treat with caution
| Transcript token | Risk | Treatment here |
|---|---|---|
| DDCP | Model name, possible mishear | Kept |
| GCV4 / OCE / O4.7 | Version or competitor code | Kept |
| MILES / MILOS / MIS / Office | Repeated in one question, not recoverable | Kept, marked unclear |
| “Learn to learn” | Mixed with “continual learning” | Listed separately, not equated |
| GDP 10% / 20% / 5% / 1% | Argumentative magnitudes, not one series | Left in their original arguments |
| 20,000 H-eq, 16,000 950s, 800B, 50k/200k cards, etc. | The class of figures the host asked not to leak | Listed as transcript figures, not cross-checked |
| “Non-compliant cards” | Ambiguous | Quoted; no legal reading |
| “One of the three bodies” | Spoken metaphor | Kept; no expansion into fiction |
15.2 Chinese–English term map
| Chinese | English used here |
|---|---|
| 愿景 | vision |
| 克制 | restraint |
| 开源 | open source |
| 合理利润 | reasonable profit |
| 十个月收回成本 | ten-month payback |
| 六倍利润 | ~6× profit |
| 持续学习 | continual learning |
| 思维链 / CoT | chain of thought (CoT) |
| Agent | agent |
| 具身智能 | embodied intelligence |
| 激活参数 | active parameters |
| 主线 | technical mainline |
| 副产物 | by-product |
| 降维打击 | advantage of operating from a higher technical plane |
| 奇点(强调渐变) | “singularity” (stated as gradual) |
| 摸奖 | lottery-style exploration |
| 正式(自上而下) | formal (top-down) work |
| 后训练 | post-training |
| 幻觉 | hallucination |
| 三体之一 | “one of the three bodies” (his wording) |
This brief does not replace the recording. Where figures, product names, or competitor references conflict with public information, the recording and any later written statement by the speaker govern.