In this episode, I’m joined by Robin Zhu, one of the sharpest observers of China’s technology and AI landscape.
We talk about some of the biggest names in Chinese AI, from Z.ai, Moonshot and DeepSeek to Alibaba, Tencent and ByteDance. But rather than just looking at who has the biggest or most talked-about models, we get into what each company is actually good at, how they’re approaching the frontier, and what Robin looks for when trying to separate real progress from the hype.
We also touch on how Chinese AI companies are operating under very different constraints than their US counterparts, yet they’ve continued to make impressive progress through techniques such as model compression and reinforcement learning. This raises a bigger question around where the value in AI ultimately sits: if models become increasingly capable, cheaper and more commoditized, who actually captures the economics?
From there, we get into the business of AI — how open-weight labs can make money, what AI monetization might look like, and whether the biggest opportunities will sit with the models themselves or with the applications, infrastructure and orchestration layers built around them.
Finally, we zoom out to the bigger picture: what China’s progress in AI could mean for geopolitics, model sovereignty and international adoption, and how investors should think about valuing these companies when the technology is moving faster than traditional financial metrics can keep up.
The AI Proem Podcast is under the AI Proem newsletter, which has over 12k followers globally. To learn more about China AI, the business of AI, and how AI is impacting businesses, please check out the newsletter here and more insightful conversations here.
Chapters
00:00 Introduction
01:06 China’s AI Race and the Rise of New AI Labs
06:39 The Compute Bottleneck — And How China Is Closing the Gap
11:37 AI Monetization: Who Captures the Value?
13:06 Why China Has So Many AI Labs — And Who Will Survive
16:10 When Does an AI Model Become “Good Enough”?
19:10 Token Rationalization, AI Harnesses and the Future of Work
25:21 Models vs. Applications: Where Will AI Value Accrue?
30:31 How Open-Weight AI Labs Can Make Money
33:07 Can Chinese AI Capture 30–35% of Global AI Revenue?
45:02 How Should Investors Value AI Companies?
49:56 Robin’s final thoughts on AI’s future in China and globally
Transcript
(AI-generated, for reference only)
Grace Shao (00:01)
Hey Robin, good to have you.
Robin (00:03)
Thanks for having me. Good to be here.
Grace Shao (00:05)
Yeah, yeah. Tell us about your coverage and your recent initiation on Z.ai and MiniMax. I think that was quite exciting. It was a huge report — 50 pages or 80 pages, was it? What made you decide that now was a good time? And what is your main takeaway there?
Robin (00:23)
Sure. Yeah, look, you know, I’ve been covering internet at Bernstein for a long time now. I’ve been covering gaming for the last number of years in Japan. Year to date, I think something like 80% of our research has been about some form of AI or other. I’ve been using Z.ai and MiniMax as the examples to effectively fill my exhibits and illustrate different points. The stocks kinda ran away from me as we were doing that. We initially thought, okay, we were gonna you know, work out what AI does or what these businesses do and then they all went vertical. there came a point in the summer I was just like, All right, you know, the stocks can do whatever they want given such small free floats and we’ll wait a little bit and the lockup expiries were coming up at that point and yeah, we picked a week Shortly after. I was in the US for a month to kinda network and do different things. and I think we got lucky on the timing, to some degree. But yeah, you know, now there’s more price discovery. There’s, you know, it seems to be a new model launching every other week. So yeah, fun times.
Grace Shao (01:34)
Yeah, sorry, we were just talking about how there’s such AI fatigue. Like, there were literally eight models over the summer and there was no summer for any of us covering AI, right?
Robin (01:44)
Are you not excited about Ox Alpha?
Grace Shao (01:47)
Everyone’s excited about Ox Alpha, but we all have different conspiracy theories, right? Well, because like I don’t wanna like you know go into these conspiracy theory holes today. Let’s focus on some of the big pictures. I do want to ask
Robin (01:58)
Okay.
Grace Shao (01:59)
You are one of the rare people who gets access to these labs and their executives. When I last spoke to you, you said you were hanging out in Beijing, meeting with some of the executives at Z.ai, MiniMax and whatnot.
Robin (02:09)
Mm-hmm. Mm-hmm.
Grace Shao (02:12)
Obviously not sharing anything sensitive, but what’s the vibe? What are the cultural differences?
Robin (02:16)
Yeah.
Grace Shao (02:17)
You know, do you think any of their personalities or cultural makeup actually, you know, differentiates them from how they go to market, how they build their products or technology, or maybe even their philosophy on AI?
Robin (02:31)
Yeah, sure. I mean, I think it’s kind of interesting, you know, I deal with investors day to day a lot. you know, the debate there is, are these labs raising prices? Are we gonna get competition? Do models get commoditized and pricing goes to you know, gets hammered and so on. you talk to the guys at these labs and it’s a very kind of singular focus on, you know, everybody thinks they’re changing the world. AGI is very much top of mind for everybody and Iterating the model is much more of a focus, obviously, compared to investors. But the vibe is very much: yeah, just keep going, keep cranking and see where we can go. Culturally, there are some quite big differences. You know, Z.ai came out of Tsinghua University. Dr. Tang is still kind of on both sides of the fence in some ways. You know, somebody else described it as being monastic. I’m not sure I’d go that far, but it is a much more kind of academic and nerdy organization. MiniMax and all the dealings I’ve had with them seem to be more commercial. You know, they’ve had a couple of pivots in terms of what the main focus has been. certainly more international than Z.ai. but yeah, Kimi’s kind of I guess in some ways halfway in between. You know, they are More international than Z.ai but yeah, you know, you’ve got kind of the more si how do I how’d you describe it? More kind of science based aspect of you know what they’re doing. So yeah, you know, these are they show through in how these companies behave, and the results that you’re seeing in terms of model progress. And yeah.
Grace Shao (04:24)
How do you think they’re defining AGI? Is it different from what SF is saying?
Robin (04:31)
What’s SF saying? It seems to be different every few weeks.
Grace Shao (04:34)
Huh.
Robin (04:35)
I don’t know if there is a single kind of monolithic, you know what — like we’re going to do AGI and it’s this thing. you know, I think the common analogy is summoning the machine god, which... But I think it’s a little bit narrower than that. I think it’s, you know, how do we get AI to iterate our models for us? How do we get into kind of, you know, I guess some definition of loose RSI or narrow RSI? I don’t tend to get into discussions about broad RSI with people, you know, where it does actually just become a little bit more religious. But yeah, I think, you know, everyone is just focused on kind of iterating the next generation of models.
Grace Shao (05:17)
And they’re all kind of facing the same issue, right? End of the day it’s compute. But potentially there’s domestic compute becoming more abundant. do you think that’s really gonna change the game here? Or is that even something that’s happening in the near term?
Robin (05:33)
Yeah, I think it is a bottleneck. Everybody is short on compute. Everybody makes comments on, you know, if you compare the FLOPs per engineer here versus in SF, there is a big difference, and it does hold back some of the progress. But despite that, you’ve seen some of the Chinese labs come up with cache-compression tricks, RL tricks to Close the gap, and I think it has been quite remarkable to see where they’ve gotten to on quite limited resources. I mean, Z.ai in particular, getting to what they’ve done with a 750B pre-train has certainly surpassed what I thought was possible without getting to a bigger model.
Grace Shao (06:22)
Okay, well then let’s take some. I wanna double click on that later for sure on what the potential implications of all that progress means later. But first, start big, high level. What’s your sense on each of the labs? Like who’s good at what? What are they each gunning for? What’s a good mental framework for us to when we’re evaluating these different labs? Because the one thing I wanna lead to the next question really is we just have so many Chinese like lab providers, I sorry, model providers right now. Like, There’s Z.ai, MiniMax, DeepSeek, let’s just call them somewhat the first tier labs. Then we have like the BAT here, I mean like ByteDance, Alibaba, Tencent. Then like randomly over the last like three to six months, we get Xiaomi, Meituan, RED, Huawei, all crowding that space. And then you have like StepFun and a few others kind of dabbling in this. Well, not dabbling, but they’re also doing this. Well, maybe considering them second tier. How do we understand this landscape and
Robin (07:18)
Ni
Grace Shao (07:18)
How do we evaluate? These labs.
Robin (07:21)
Yeah, I think the way that we kind of framed it and when we launched coverage was, you know, I think there are three frontier labs if you look at the latest models. The concept of the Pareto frontier is quite important because there isn’t sort of a monolithic kind of first place. You either have to be the smartest model at your price point or vice versa, the cheapest model for a certain level of intelligence, however you define that. So you can have different spots on the frontier. For example, when Kimi K3 came out, it obviously was smarter than the latest GLM, but it was also something like three times as expensive. So I happen to use both in my day-to-day. and they’re not kind of direct substitutes immediately to each other. but let’s say, you know, we said and I stand by this view that there are three frontier labs: Z.ai, Kimi/Moonshot and DeepSeek. I think amongst themselves, the consensus is that Z.ai is good at post-training, Kimi has the biggest pre-training scale and is good at pre-training, and DeepSeek has this crazy infra capability, so they can run super-high tokens per second and so on. So, you know, they are good at different things based on their different backgrounds. The internet companies — look, Alibaba’s probably spent the Most time and effort to try and develop a frontier suite of models. You know, Qwen is frontier-ish. It’s kind of up there. And then Tencent’s come back really into the conversation in the last six months with Hunyuan 3. I think initially people ignored the preview and then more recently it’s become clear that, you know, the machine that builds the machine is now working and they’re iterating towards Hunyuan 4 sometime this year. So you know that’s that.
Grace Shao (09:13)
I think Hunyuan 4 preview is coming out this Friday actually. It is. Somebody just told me. I think that’s public information soon
Robin (09:17)
Is that right? can I quote you on that? Yeah. Cool.
Grace Shao (09:24)
Or now. Anyway, yes, go on.
Robin (09:26)
It is now. Yeah. well, You know, they said it’s coming this year. I kind of assumed it was coming in Q4, but cool. That sounds encouraging. ByteDance is funny, right? Because they kinda raced out into a big lead with Doubao. It was, you know, it was considered to be the — well, it’s, I guess it still is the chatbot app in you know, compared to Qwen or compared to Yuanbao. And then I think they ran into the issue of, well, okay, we’ve grabbed the users that we can grab, and then when they tried to monetize it through subs, the paying ratios were very low. and so now I think there’s been a bit of a pivot towards, do we do enterprise? My understanding is they’re gonna start doing ads within Doubao in the second half of the year. so there’s a couple of ways that they’re going through. But yeah, I would say that Tencent and ByteDance are more focused on applying AI through their ecosystems. Alibaba’s much more focused on
Grace Shao (10:22)
But they want to use their own models, Right? They want to use their own models. So it’s like without really strong models, how do they apply? And this brings me, sorry, I’m just hijacking this whole conversation, but it brings me back
Robin (10:30)
Mm.
Grace Shao (10:31)
To like Ben Thompson’s recent interview with Patrick Shonas. He was really interesting. He just goes on about, you know, end of the day, it’s advertising and consumers will not pay. So it seems like, you know, all these AI application companies will just become like they will just monetize the same way the internet companies monetize. And then it’s
Robin (10:45)
Everybody sells ads in the end, C Yeah, I think that’s I think that’s A plausible yeah, I think that’s a plausible end outcome. where I would differ a little bit is something like a WorkBuddy, where people are just paying for effectively tokens and task completion. but yes, on the kind of base consumer layer then yeah, advertising, monetizing merchants who Effectively pay for access to people doing stuff on WeChat or elsewhere ends up being you know, Douyin started to monetize some of the local service recommendations recently. again through a kind of quasi ads model. So yeah, I think we do move in that direction medium term.
Grace Shao (11:27)
Then what about all the rest like I just talked about? I think like 10 other companies are serving up models these days. How do you make a or actually let me re-reshift the question? How why
Robin (11:34)
I tend to think of those as
Grace Shao (11:40)
Let me reframe the question: why are there so many? Should we be expecting some calls consolidation? Like, because I don’t expect you to comment on every single one of them, how they’re different. But the fact that we just have like 20 different labs
Robin (11:50)
Mm-hmm.
Grace Shao (11:52)
Offering models at this point. It doesn’t seem very economical, but and also aren’t they all just fighting for the same compute, same talent at this point? How do you view all Of that?
Robin (12:00)
Yes, there’s only so many PhDs that you can hire out of these places. Look, I think there will be fewer players at the frontier or frontier-ish. you know, I’ll pick on Meituan since it’s a company I cover. But they came out with LongCat-2.0. I think initially there was some excitement, and then people realized that it’s got the reasoning capabilities of Hunyuan 3, which has got a quarter as many parameters or something. So or less. and I do think that over time, you know, if you think about what’s important here, it’s access to a data pipeline, it’s the ability to turn that into RL environments and then train these models. and I do assume that the you know, the labs and the very biggest internet companies will kind of, you know, be thereabouts. Some of these smaller names that you’ve just mentioned, I think you know, in the Meituan example, like what’s the difference between having your own model versus using DeepSeek or something? Like, just You know, fine-tuning DeepSeek or something. I do think there’s a kind of open debate there. and you know, certainly costing them a decent amount if you look at their accounts. So I do think there will be fewer players at the frontier. You know, I think
Grace Shao (13:14)
So
Robin (13:14)
If you want to have some kind of basic search engine that is AI powered, then so be it. Do you need a complicated long horizon agentic model to underpin every internet platform? No.
Grace Shao (13:27)
What is driving all this kind of effort to even build their own models? Just FOMO.
Robin (13:33)
I think it’s partly, “we want our own thing.”. I think it’s innately because there’s so much competition in the internet space that there is this kind of insecurity around you know, if we don’t have a model then do our peers who do then find a way to get an upper hand somehow. And so we need to at least understand the technology. That bit I get. I just don’t think that translates into long term commitment into, you know, training ever larger models infinitely.
Grace Shao (14:02)
Why don’t we see that kind of phenomenon in the US as much? Like the internet players aren’t just all rolling out their own models.
Robin (14:11)
They’re all too busy serving compute to the two guys at the front. Seems to be what’s happening. Yeah. I think generally
Grace Shao (14:16)
All right, let’s
Robin (14:19)
Generally, there’s a lot more duplication in China than
Grace Shao (14:24)
Mm.
Robin (14:24)
In the US where the you know these players kinda keep to their own lanes a bit more.
Grace Shao (14:29)
Yeah, yeah. Well, on that topic, this is like relevant, but much of a conversation around AIs that these models are being commoditized, but you do have a more nuanced view. reading your recent report, you know, you’re saying
Robin (14:40)
Mm-hmm.
Grace Shao (14:40)
That, you know, frontier models versus good enough models kind of have different value within the long term ecosystem. At what point do you think a model becomes good enough that the user simply stops caring whether another model is theoretically more intelligent?
Robin (14:55)
Yeah.
Grace Shao (14:55)
Do you think it’s use case dependent or you know? Client-, customer-, end-use-dependent or how do we understand that?
Robin (15:04)
Yeah, I think initially, you know, agentic AI kind of exploded in Q1, and then everybody got super excited and you know went both feet in to try and work out what they could do with it, hence the token maxing movement. and then subsequently I think as especially as we started using agentic AI more and more, the framework that I kind of zoomed in on was more around User perception, right? And you can define that however you want in terms of use cases, in terms of you know who the user is. And I guess the point that we tried to make was instead of there being some kind of objective standard by which AI models become good enough, they become you know, AI becomes good enough when you can solve a task that you’re trying to solve. And If you then have more than one AI model or if you have multiple AI models able to solve the same task, then the arbiter of who you give the task to then moves away from reasoning capabilities because at that point, you know, by default, if multiple models can solve the problem, then you move on to cost and availability and you know, in some cases UI/UX and your kind of taste of one model versus the other in some cases. So yeah, I you know, I think If you kind of lead on to that, then we do think in a couple of the companies have talked about this in different ways, is that you will have a frontier where people will pay higher and higher ARPUs for more and more specialized and longer and longer horizon task completion. you know, you go from ordering bubble tea to doing agentic commerce to doing, you know, more serious work stuff to frontier science. With the number of tasks that you solve probably get fewer and fewer and the ARPUs just scale infinitely almost. And meanwhile you’ve got this kind of behind the frontier bit where yeah you can solve a task with a good enough model. you know, we were in discussions with or talking to execs at Tencent who said that on WorkBuddy there’s probably thirty percent you know performance-driven tokens and versus seventy percent what I think was said was value-driven tokens where it’s like, you know, you can solve the task with A cheaper model. the way that Z.ai tries to frame it is, you know, you have thirty percent and then maybe fifty percent, which you can monetize and then the last twenty percent is just given away for free on Doubao. but yeah, you know, you have a split in the market as a result of you know, using models to solve what it can versus solving problems that you know multiple models can do versus the really simple stuff.
Grace Shao (17:50)
And is that the responsibility of the, I guess, the provider of the tokens or the user? As in, should I be routing my own models left, right, center to optimize my cost or what? Or should the platform be
Robin (18:02)
Yeah.
Grace Shao (18:03)
Doing that for me?
Robin (18:04)
I think people ran into that themselves because I mean the irony for me is that Fable being as expensive as it was was probably the catalyst for people to realize, you know, maybe we shouldn’t ask Fable for the weather because it costs you five dollars to do that or something. and and then at the same time, I think in June, July there was this explosion of open-source Chinese models that you know meant that the number of alternatives Then expanded and people started to kind of think about, maybe I should rationalize or, you know, at least match the value of whatever task I’m solving with the cost of the tokens that you know, that I’m spending. In the first half there were these crazy kind of anecdotes like folks working for the internet companies having quotas of upwards of like a thousand dollars a month and, you know, me Saying to them like I know what you do, you make slides for your boss. You don’t need a thousand dollars of tokens a month. and eventually I think the internet companies backed away from that. So yeah, you know I think we’re seeing that kind of rationalization movement happen.
Grace Shao (19:13)
Yeah, yeah, there’s definitely been a cap I’ve heard for especially what they call like people who are just working with documents, definitely definitely don’t need to be token maxing at all. I am an advocate for not using AI for simple things. Like when you’re asking about the weather, maybe you should just turn on your weather app, like honestly, and not like burn compute
Robin (19:29)
Yeah.
Grace Shao (19:30)
On that. Anyway, you use the term token rationalization in your recent report. Tell us about that, because you kind of alluded to it already.
Robin (19:39)
Yeah, yeah, yeah. So I think it’s, you know, again, it’s this idea of matching the value of what you’re doing versus the cost of the tokens that you’re spending. and it kind of ties back to this Pareto frontier where you know, you have different levels of complexity, you have different levels of cost associated with different models based on their scale and other factors. And you know, you try and be smart about You know, how much you spend, I guess. Like one of the kind of ways I’ve settled in is to have worker bots. I run my own Hermes setup with different bots and I have my worker bots that are relatively cheap and do what they do. and then I have a checker of homework at the end that makes sure that nothing stupid happens and they red team each other and so on. So I think, users kind of find their own ways to do that. Maybe You know, outside of the kind of pro users, then there is the role of the harness or of the kind of you know, the WorkBuddy-like app that then decides for you, this goes to this model, that goes to that model. That makes the overall experience more optimized.
Grace Shao (20:49)
Why don’t we just jump it straight into how you use AI for research? Like I wanna hear more about it. Like, how do you use Hermes? How do you run your own models? Like it’s not a Very
Robin (21:00)
Yeah.
Grace Shao (21:00)
Common sell-side analyst approach. I think you’re definitely like AI-pilled to the max out of all the cell
Robin (21:06)
I am
Grace Shao (21:07)
Sell-side analysts I speak to. Why don’t we talk about that first before
Robin (21:12)
Yeah.
Grace Shao (21:12)
I get more into the analysis? Like I’m curious.
Robin (21:15)
God, how long do you have? Look, I started with OpenClaw, round about the same time as when everybody got excited about OpenClaw. And then I very quickly fell out of love with it and then almost by accident happened upon Hermes around the same time. I use it for a number of things now. I use it for kind of work research as a way to gather information. these bots have a way of scraping data from websites. That I can’t manually. there was one instance of a website where you know the website would show you the last 12 months of data and then the bot went inside and after a while said, here’s an API that allows me to download the last 10 years of data, which is great. maybe it’s less great for the guy operating the website. But it there are things like that, or I can now I’ve set up bots to monitor Twitter and Reddit sentiment when a new game comes out, or you know, I’ve built You know, reasoning loops to try to forecast game sales, look at or actually one of the biggest time savers of all has probably been the ability to summarize podcasts. Like in the old days, pre AI somebody sends you a two and a half hour podcast, you’re like, great. Like I’m sure this is super interesting, but I’ve just lost my Saturday morning. Whereas now I can basically have the Hermes bot gonna summarize it, give me key quotes, key insights, timestamps, you know I probably then go back and listen to what was actually said on maybe 10–20% of the two and a half hours. So yeah, it’s been, you know, a good time saver. The other thing is you know, just the ability to knock around ideas on my phone while I’m walking around. Like there was one time when, you know, my wife and I took our daughter to some tourist spot that I’d been A lot of times and they were kind of going around sightseeing. I was behind them having a chat with my Hermes bot and by the end of the four hours of walking around I developed most of a note to write down. So yeah, it’s it’s you know, there’s different ways that I’ve tried to use it. I’m sure we’ll find more over time.
Grace Shao (23:26)
So you just exposed yourself For not actually listening to my AI Proem podcasts when you do say you listen to them. You’re probably just getting Hermes to summarize them for you.
Robin (23:36)
No, I listened to you.
Grace Shao (23:38)
All right. You listen to this episode.
Robin (23:39)
Hahaha.
Grace Shao (23:41)
Okay, let’s get back to the serious stuff. Okay. I think I think that’s really interesting because I think someone who actually uses it, like understands it differently from just pure observer. But do you think then, just going back to the last conversation, but do you think then frontier model providers still capture the most of the economics? Or do you think the value is gonna move to applications, orchestration layers? Because early in the conversation, you know, we were talking about like, you know, people are like more mindful of the costs now. but this is how these frontier labs make money. So doesn’t it then challenge their existing business model?
Robin (24:15)
Yeah, I mean there is this ongoing debate about value capture, you know, between the semis industry where all the stocks have gone vertical this year and then less so in the last couple of months. and then the hyperscaler layer, I mean that you know, you if you look at the US internet companies, spending AI capex is alternately good and bad every other few months, you know, depending on what gets reported. And then the labs themselves, and then, you know, increasingly you’ve had these kind of harness type debates. The one that has struck me as being very interesting lately is that the competition between first and third party harnesses. If you’re an AI lab, then you know the harness is something that basically puts — if you think of the model as being an answering machine or a reasoning machine, the harness actually adds persistent memory, reference files,, you know, tool calls and The ability to do stuff on a kind of recurring and ongoing basis next to the model. And that’s actually what makes it so for example, the Hermes construct is what makes these bots be able to work with you much more like a human worker can. Right. And so there’s been the debate of, well, you know, it’s strategically necessary for every AI lab to have its own harness, whether it’s Claude Code, whether it’s Codex, whether it’s Z Code and Kimi Code and so on. Or do you just end up with kind of WorkBuddy and that acts as an orchestrator across multiple models? or do the first party models then allow other models into their own kind of space? So that’s a debate I think is not going to be solved anytime soon. I think it’s gonna be useful to kind of see how it plays out. but at the end of the day, especially in the Chinese context, I do think that, you know, if the strategic need is for Frontier reasoning capabilities, then on some level these labs have to survive. Right. If you assume that essentially they get zero part of the value, then they will cease to function or they will stop being able to fund themselves and fund the next training run. So I think they will have to retain some of the value in order to keep doing that. and if you think about which layer of the Chi of the tech stack in China that’s most likely to get overbuilt, it’s probably the compute layer. Which also argues in favor of the labs getting, you know, some of the value. So I think it’s yeah, I think it’s an ongoing debate.
Grace Shao (26:41)
Interesting, you just brought up actually like a lot of these labs will have to create their own products to actually still capture some of the value beyond just infrastructure layer, right? But actually, I think Sam Altman was just on a podcast like a few days ago. He was talking about how he’s like actually bringing everything back. they’re cutting products, right? Like no longer doing a bunch of different products. They’re saying that they’re a platform business instead they only want to use ChatGPT’s interface, they’re even like renaming Codex or something. Like, do you think that is the future for like the kimmies and the mini sac max of the world ‘cause then my argument or my challenge against that is then how could they compete with the big tech and channel where they have just like such broad reach, I guess. They’re like with these super apps and everything.
Robin (27:27)
Yeah, so that’s that’s the interesting part where if you’re an AI lab, you know, right now you’re using, you know, some of these third party harnesses as distribution. over time, you know, do you have to retain some parts of the reasoning capabilities like, you know, for example, cyber is one of these niche things that you can do with an AI model. Does that need to go into a you know, power user grade or consumer grade AI harness? Or, you know, there are different Flavors of model that do different things, not you know, maybe you don’t put all of them onto the generic kind of third party harness. but that is something to work out. I mean in Sam’s case, I think there’s a kind of semantic thing here where you can name it what you want. to me the accumulation of personal context behind the model through repeated engagement with it and it kind of learning what you do and You training it to do different things as skills, setting up a network of expert tool calls that you can kind of call on. Like that stuff is actually in my mind, where a lot of the long term value lies, or where the usefulness of the AI kind of goes and resides in the you know, as you use it repeatedly. So yeah, and you know, there are a few ways where I think this can play out. One is potentially, you know, there is obviously the The version of the world where everybody goes on to a third party harness within you know that’s run by one of the big internet companies. The alternative is that, you know, if you are a more pro user, more specialized user, maybe you do need some of the specialized functionality that, you know, the AI labs keep to themselves. and then you have a range of these outcomes. yeah, very complicated question. I don’t know if I have a fully formed answer at this point.
Grace Shao (29:18)
Fair enough. but let’s look at monetization in general for the open-weight labs. You know, obviously I think from the Western perspective, it’s like the biggest question
Robin (29:25)
Mm-hmm.
Grace Shao (29:25)
Is always like how do they make money? How do open-source models make money? How do you see like the current API sales? Is that just a durable revenue pool? or is that not gonna be enough to sustain the capital they need to continue to train?
Robin (29:43)
Yeah, I mean right now, you know, back to your earlier point about compute constraints. you’ve trained these models, there’s been this explosion of interest in them. I think one of the biggest constraints that they have faced is the lack of silicon is constraining their ability to serve up APIs. You know, I feel this pain every morning Asia time where, you know, I ask a Chinese model to do something at nine AM Hong Kong time and it’s rate limited all the time because everybody else is trying to do that at the same time. And as that gets resolved, presumably, you know, that induces some demand and your ARPU goes up. yeah, I and there are probably lumps around new model releases and whatnot that means that you know it’s spiky rather than that being this smooth curve upwards. But yeah, I do assume that expands over time. You know, all of these companies distribute via the internet platforms as well through, you know, ModelScope and workbody and You know, other things. and then I think to me the interesting thing that’s kind of happened recently is this idea that they are now starting to try and charge or take rate from the global inference providers, right? Like Kimi has gone and signed these deals with different inference providers, essentially charging them you know, you can call it a royalty fee or a or it’s some kind of licensing agreement. But essentially altering the license so that if you are Looking at the weights for academic use or whatever, then that’s fine. If you’re using it to actually just host it and make money, then you have to pay them something, which I think makes sense. So that’s a route that they can kind of pursue to try and capture some of the global economics.
Grace Shao (31:26)
Yeah. So going towards like a bit more controlled commercial economics. Okay, you estimated Chinese models to be able to address roughly thirty to thirty-five percent of global AI revenue, despite the current headwinds with geopolitics and whatnot. Walk us through that thinking. Thirty to thirty-five percent is quite a large pie. I think even a year ago when I spoke to some of the labs, they were jokingly saying, even if we get five percent, that’s enough money for our business, you know.
Robin (31:54)
Yeah.
Grace Shao (31:55)
I mean, but that’s one lab. I guess cumulatively it’s it’s it adds up. So tell us about your thinking on that.
Robin (32:01)
Sure. I mean it was more of a top down kind of estimate based on what was Attainable or what was kind of in a you know in the context of geopolitical realities what was realistically kind of accessible to these labs. And you know we had made these TAM estimates by region. I think US was like half of global or you know maybe even a little bit more than that. And then China obviously we assumed was a captive market. We ended up assuming a very, very limited access to the US market because of the geo issues. and you know The thing that there’s been a few things that have happened since, like, you know, potentially well, Ramp, I think at one point said that the like f you know, five percent of their highest engagement AI users were playing around with AI models from the Chinese labs. and then you had Microsoft that was allegedly thinking about using different AI models from the Chinese labs to power copilot. So, you know, have we been Conservative, have we been kind of you know, is it that there truly is no access to the US market? Question mark. But you know, in Europe you know, we’ve assumed some access, you know, not unconstrained access. I guess if you’re Airbus, you’re probably never gonna use a Chinese model for obvious reasons. But then you’ve also had you know Mistral Mistral’s kind of turned itself from being a you know frontier lab to something that now helps Z.ai to distribute GLM. And so, you know, there are these kind of future permutations I think are gonna be interesting. That means that, you know, that Europe is accessible to the Chinese AI labs to an extent. and then in the rest of the world, I mean you’ve got, you Z.ai, for example, Z.ai doing sovereign projects with Malaysia, with the Middle East. that presumably then acts as a bit of a kind of beachhead for them to go and do other stuff within these markets. So Yeah, you know, we assume more access to these other markets where, you know, the geopolitical picture is probably you know, more more kind of open to the Chinese labs. So it’s a top down estimate. You know, that it doesn’t mean we think the Chinese AI labs will have thirty, thirty five percent market share. You’re still kind of contesting these markets with OpenAI, Anthropic and others. but yeah, it was more of a kind of, you know, how much of the TAM is actually open to you?
Grace Shao (34:24)
And let’s just say like geopolitics side, like you mentioned, like obviously government agencies are not going to use Chinese labs, but like a lot of companies might, like, you know, companies with less like strict compliance on this, then what does continued progress with China’s open-source models mean for like what would it mean for these US frontier labs, especially as of now realistically, there’s really two labs left at the kind of frontier really fighting it out. and they’re not you’re not seeing them lowering their Cost or opening up their weights.
Robin (34:53)
You’re gonna get hate mail from Elon Musk after you upload this.
Grace Shao (34:59)
Well, if he watches this, it’ll be great.
Robin (35:03)
No, look, I think you know, I think there’s going to be a a mix of model use in in the market, right? Like isn’t something I think a lot of people have underweighted is what Alex Karp has been saying, where, you know, companies need sovereignty over their own data, you need ownership of what you’re doing. He’s obviously talking about his book, but you know, I do think that there will be a variety of solutions. you see, you know, folks from Databricks and DoorDash posting about testing Chinese models on Twitter and I think it’s interesting that’s, I think that will continue. I think you know these companies will find ways to orchestrate across different models. And the fact that certainly compared to the US labs, these are generally smaller models. if you can get most of the reasoning capability from for much lower token costs, then yeah, I do I do think that you know in a growing section of Use cases that you need, that these will be good enough. Like within the home market, obviously they fight to be frontier. but outside of China in the global market, they are that kind of, you know, eighty percent cheaper for or, you know, much cheaper for most of the capability kind of market positioning.
Grace Shao (36:21)
All right, let’s zoom in on the companies themselves. I want to kind of touch on the labs, especially the two companies you just wrote about in your initiation report, and then we can talk about your long-term coverage of ATs. Just start with what’s your bull case on
Robin (36:33)
Okay.
Grace Shao (36:34)
Z.ai? It’s no secret that you love them. Why? Well, what’s your thinking on that? Like, do you think they would
Robin (36:43)
Yeah.
Grace Shao (36:43)
Just be like the leading research engine research lab in China?
Robin (36:48)
Yeah.
Grace Shao (36:49)
Frankly, not nationalized the same way that DeepSeek is likely more to commercialize. Like, I don’t know, but you just raise your eyebrows. Maybe I’m understanding that incorrectly. Help us understand. What do you think of Z.ai these days?
Robin (37:00)
Yeah, look, I’ll save the DeepSeek comment to the last. But you know, I do like their ability to iterate these models. You know, they came out of Tsinghua University and I think that relationship helps on some level when it comes to expert domain training data. You know, when you talk to them they emphasize that they have this data advantage, which I think you’ve seen through some of the RL progress that they’ve shown. and, you know, the ability to
Grace Shao (37:27)
Sorry, what Is their data advantage? What is their data advantage? They’ve said that
Robin (37:30)
As an
Grace Shao (37:31)
To me too, but I don’t know what that means.
Robin (37:33)
Yeah, I think it’s, you know, if you can buy data and you can, you know, acquire data from experts, you are effectively paying people to write down what they know. but there is also the process of turning that into verifiable you know, like RL environments where you have verifiable kind of end goals or checkpoints that the model needs to hit, or how do you verify correctness or not? And Turn it into something that’s a lot more structured and you can feed it into the RL pipeline to actually train models with it rather than just you know, I sit there and write a hundred page thing on how to do equity research or something. Like I you know, there it has to be there yeah, there has to be kind of reasoning gates and a way to kind of let the model kind of as assimilate that information. so you know, the I think in the GLM-5.3 Release paper is actually really interesting in the sense that they emphasize look RL is all we did and effectively they’ve taken you know data and you know translated that into different RL environments and then they’ve used that to try and iterate the model in different ways. So you know I do think that’s a useful skill to have. The fact that they have a seven fifty B model that’s as good as it is is interesting to me. The fact that, you know I think this is known. I mean that the next big boy model is coming later in the year, you know, call it October, maybe a little bit earlier, maybe a little bit later. But at which point you start having probably the best seven fifty B class model, or certainly it to me it is at the moment. and if you have something that’s competitive in a much bigger model size, then you actually occupy two parts of the of the Pareto Frontier front potentially, which is Interesting strategically. so I but I think most of all it’s just that, you there are these three labs I think are frontier. one of them is a little bit captured in terms of, you know, having to answer to the government to some degree. Kimi I like as well. It’s you know, I think Kimi’s doing some really interesting things on a number of fronts.
Grace Shao (39:50)
Yeah, I was just gonna say actually, like Kimi obviously you don’t cover officially now given that they’re not public yet, but you know, they kind of reset expectations around CI. I think when five point two came out, they felt like the world felt like CI was the leading l lab coming out of China. Kimi K three kind of put themselves on the global stage again. In fact, I think they did a you know, really, really big marketing splash globally, and captured a lot of attention. And then given that they don’t have the kind of CAI ent like Was it entity list complication. They actually
Robin (40:23)
Yeah.
Grace Shao (40:24)
Have it easier with international expansion, but it just kind of looking at, you know, Kimi K3, how do you view Moonshot and theirs their positioning right now?
Robin (40:35)
Yeah, I mean the fact is that they have the biggest Chinese pre-train, right? And the Kimi K3 is a very capable model. it’s significantly more expensive, but at the same time, you know, if you’re among c corporate customers, I think there is the argument to say that you just you know, it’s cheaper than the US labs anyway and you just pay for the best capabilities on some level. But yeah, look, I think, I think they will be up there in the fullness of time. They will hopefully get listed before too long. You know, the last time I asked them the answer was soon. so we’ll see. But yeah, I I think they’re you know, what happened with GLM-5.2 and K three was was kind of interesting because there was a point before the summer where we could have launched coverage on these on these AI labs and almost Around the same time basically GLM-5.2 happened. It was great from the perspective of somebody that used these models, but then I think Z.ai’s share price went up like fifty five percent in the week or that week or something. And it like, All right, maybe we’ll take a break and see how see what see how it goes. And then it got to the point where I think, you know, Z.ai’s share price was pricing in being a winner-take-all type winner, at which point, you know, when Kimi K3 came out then there was an unwind of that expectation. I would argue that there shouldn’t have been that expectation, but you know, go figure. So I think now, you know, I still think that these are the two to watch. DeepSeek obviously will always be up there, but I personally find the Kimi and GLM models way easier to use.
Grace Shao (42:18)
And Why was it that, you know, when Z.ai and MiniMax went public that it felt like MiniMax was more of the market darling, or at least investors in Hong Kong were buzzier around them.
Robin (42:29)
Yeah, I think there were two reasons. I think one was MiniMax was considered to be a lot more international. or it was it was much more international. And then the other thing was given the Entity List listing that Z.ai had had, that it was kind of thought that they would find it more difficult to expand internationally. And the other the other problem that investors had with Z.ai was the on prem Segment, which, you know, was kind of this it reminds people of the bad old days of China Enterprise Software, right? Where, you know, you had these companies kind of toil for years and years without really getting anywhere. and so that was, I think, the initial impressions. we put out something in quite early on, I think after CNY, where you know we took a deep look at these companies. I think I was always of the view that You know, model reasoning capabilities are more important and the state of these companies today versus six months ago versus six months in the future is gonna be so wildly different that it’s yeah, you need to evaluate the machine that makes the machine more than kinda where they are at any given moment, which I think was you know, people were guilty of in January.
Grace Shao (43:41)
And you make the somewhat sacrilegious sell-side argument that traditional financial analysis of these labs can be almost irrelevant essentially. Like are we effectively valuing these companies right now? Like what do you think we should be actually looking at when we are putting valuations on these companies right now?
Robin (43:56)
Yeah, I think the market really struggles to value these companies because, you know, everybody agrees that AI is a big deal and you can have these debates on how many trillion dollars of TAM AIs going to be in the fullness of time. I’ve largely given up having that conversation. It’s just it’s going to be big. and you know, investors do value these stocks on the basis of multi-year ARR trajectories and you know, revenues multiple years out. But yeah, like, you know, these companies are about to report first-half earnings. we’re about to I mean, there are things that you can watch out for, like inference margins and you know, obviously the revenue growth and How AR converts to revenue and so on. But you know, f for example in the case of Z.ai, like you’re you’re basically staring at a bunch of numbers that reflect GLM-5, GLM-5.1, which is ancient history in AI terms. So it’s kind of useful, but not really at the same time. yeah. So to me it’s much more important to just kind of look at the iteration, look at the architecture tricks that they are coming up with and the ability to kind of, you know, scale these new innovations to much bigger models. And you know, if you can do X then you should be able to do Y, and then what does that unlock in terms of capabilities? so yeah, I do think it’s you know, I do think that Stuff is at least as important as kind of scrutinizing the numbers. Even if you know it’s kind of my job to multiply two numbers together at the end of the day.
Grace Shao (45:31)
It’s more important to be looking forward than kind of looking back. but like then I have a question that’s like, you know, StepFun has already
Robin (45:37)
Mm-hmm.
Grace Shao (45:37)
Filed for the IPO. we just said Kimi is likely going to go public in Hong Kong too somewhat, sometime this year, next year, whenever.
Robin (45:44)
So yeah.
Grace Shao (45:46)
We’re looking at like four leading labs already, just the Hong Kong Stock Exchange. Then we have obviously
Robin (45:51)
Mm-hmm.
Grace Shao (45:51)
The BATs, which I want to talk about later as well. Like, how do we understand? Is this not a pretty crowded space? Like there’s a lot of labs going public in China.
Robin (45:59)
Okay.
Grace Shao (45:59)
Does that make sense? How do we understand that? How do we pick the winner? Robin, no one How do you pick the right stock?
Robin (46:07)
How do we push the right? I’ll push back and say that this is the least crowded new-tech cycle in China that we’ve seen so far. you look at you know in the past, you know, EVs and batteries and you know, or what’s going on with humanoid robotics at the moment. you know, the AI labs piece has probably been one of the more concentrated fields that we’ve seen. And within the names that you’ve Mentioned, I do think that there are kind of there’s a clear hierarchy of, you know, which ones are closer to the frontier. I think DeepSeek has decided it wants to embrace the national champion role and potentially list in the A-share market, which is fine. But yeah, I think, you know, I’ve I’ve said for a while I think Z.ai and Kimi are, you know, my picks for the frontier. I think, you know, based on what they’ve done, based on the, you know, I the architectural innovations, based on the adoption of their innovations by other labs is kind of one thing that I watch for as is being quite telling. yeah, I you know, I guess there will be more than just two players in the Hong Kong market. But yeah, I think my view is, you know, like in the US where you’ve seen fewer players at the frontier over time, I think that will show through here as well.
Grace Shao (47:37)
All right, let’s move up the stack. while you’re covering BAT, you’ve been covering them, well, Tencent and Alibaba
Robin (47:42)
Mm.
Grace Shao (47:43)
For quite a while, just given that byte dents is not public. you know, what is your mental model thinking through these big techs in China right now? Clearly they have a bit of FOMO, they don’t wanna be left behind, they don’t want to just be known as their internet as the internet phase. They’re all in IAI. They have the benefits
Robin (47:59)
Mm-hmm.
Grace Shao (47:59)
Of talent, they’re the benefits of money, but somehow the like just like the US, they’re not the ones actually pushing
Robin (48:04)
Yeah.
Grace Shao (48:06)
The frontier. How do you think we should think about that?
Robin (48:10)
Yeah, I mean psychologically, they are quite different businesses. Like, when you talk to Tencent, the focus is a lot more around the application layer and how do you apply AI and agentic functionality within ecosystems like WeChat, within, you know, WorkBuddy is potentially a new platform for them, you know, on the AI front. you’ve got the games and ads businesses, which are I actually think are good platforms on which to apply AI and generate kind of benefits. But you know, it’s much more focused on the application layer. And there was a comment in the latest slides from the Q2 earnings that said, If all else fails, then we’ll just rent out the compute to whoever else. so that’s that’s considered to be
Grace Shao (48:53)
I’m dead. I love how candid they are.
Robin (48:57)
So that’s kind of their psychology around AI. Alibaba obviously has you know you’ve got the e-commerce business, which is kind of stuck in this retail growth environment that’s not really growing. And the cloud business is you know has — well, you know, two years ago it was growing like plus eight, now it’s growing plus fifty, in the upcoming quarter, give or take. and so Yeah, it’s become the new thing, right? Where, you know, the hope is that they have Qwen, the hyperscaler layer, and T-Head, which is one of China’s better ASIC programs, which is worth something. And so, you know, to try and integrate that as a stack. But yeah, you know, if you look at the kind of growth algorithm of the business itself, it’s probably the compute layer that’s driving a lot of it. Yeah, just in terms of renting out capacity. So yeah, these are these are quite different businesses.
Grace Shao (49:56)
Yeah, and But the thing is they’re still going ahead, like you said, there’s Qwen, there’s Hunyuan, there’s Seed. should they still be in this model game, in this in this extremely competitive game, or do you think they should be
Robin (50:07)
It’s
Grace Shao (50:08)
Focusing on, like you said, like just plug in all the other models, focus on growing their existing business, right? Like how do you make up it? How should
Robin (50:18)
Yeah. I mean I
Grace Shao (50:20)
They balance that?
Robin (50:21)
Mean cloud and AIs probably Alibaba’s core business at this point. Like the e-commerce is kind of the cash cow that funds everything. but this is clearly the future and you know they’ve guided explicitly for cloud growth to be more than forty percent for the next bunch of years. so to Alibaba this is the core. in Tencent’s case I think there have been There have been multiple debates, you know, that I’ve had with investors around, you know, do they need a super frontier model? Do they need the best model in the market? Or do they you know can they just be the orchestrator layer? Can they just be WorkBuddy? Can they just, you know, use games or ads as a way to kind of monetize AI? I think right now the approach seems to be let’s do everything and see what sticks. or you know, like hopefully everything sticks, but you know, that’s That’s still the strategy and the you know the I guess one benefit that Tencent has is that they do generate a lot more operating cash flow through the core business that then funds a much bigger bonfire of capex over the next whatever number of years that allows gives them some level of optionality.
Grace Shao (51:30)
So obviously a lot of that money is also going to building out compute right now, right? So do you think China’s compute build out could be a double-edged sword at one point? It might erode some of that scarcity on pricing power and inference. Or, you know, some people are writing about overcapacity on compute. Like, is that a thing?
Robin (51:53)
Yeah. I mean, look in the in the in the infinite long run, and if you just kinda take that Tencent comment of if all else will rent out the compute, like if everybody builds compute with that as the fallback option, then the reasonable terminal outcome is that you get a big overbuild of compute, right? Just logically. but so that is a concern. And in the long run, you know, I guess you can make the argument that every new tech Cycle in China has ended up in some kind of overbuilding in the end. to me that’s most likely probably to happen in the compute layer because the level of specialization and tech and you know differentiation that’s required in semis is reasonably high. In AI labs, if you really wanted to be frontier, then there’s a level of math and science capability that you need, whereas standing up boxes with servers in them feels less Complicated. you and I probably couldn’t do it, but with enough money and help you know it should be doable for large corporations. so yeah, I do worry about that. I mean, like it’ll probably take a long time because you know even growing 100% a year, it’ll take a while for the domestic semis industry to catch up with demand. But you know, for now, I think compute tightness and You know, cost pressure in the supply chain is probably you know, it’s almost a good thing in the sense that it reduces the risk of everything kind of collapsing on itself and pricing you know, price wars and things of that nature happening in AI in China.
Grace Shao (53:34)
But compute abundance will be good for consumers, right? Well, at least for end users. Which is not a bad thing.
Robin (53:39)
Well yeah, so yeah. Well, I think there will be the, I guess obviously initially when there’s more compute, yes, you know, everybody has more capability to serve up more inference demand. And so the industry grows. I think and then obviously, you know, the hope is that there is a balance between demand and supply. In practice that almost never happens. and when you end up with an environment where there’s a you know, there’s a number of good enough models, there’s a lot of compute. Then it’s probably more likely that you know on one hand, you know, the cost of compute then goes down, but yeah, there’s more likely to be price competition in AI inference at that point than today.
Grace Shao (54:19)
And we start another round of juan. There’s never-ending juan so you talked a lot about you know how much the harness around a model can change the actual output. And I think it’s quite topical around the big tech right now. do you think we’re putting too much focus on how smart the base model is? do you think eventually really like the focus should be on memory, tools, routing, verification?
Robin (54:42)
Yes.
Grace Shao (54:44)
And then therefore like these big tech companies actually have A lot more experience in building these products, understanding consumer user behavior,
Robin (54:53)
I think both layers matter. I think it’s you know, if you’re gonna reduce the question to extremes, then if you just have the orchestration layer and you don’t have AI, you know, the reasoning capabilities of the model, then that doesn’t work. And I think the vice versa also doesn’t work in the sense that, you know, then you kinda kneecap the model’s ability to complete tasks and so on. So I think you’ll see the labs and The big internet companies all try and compete on both layers. one you know, back to your earlier question of, you know, why should all of these companies be developing AI models? one of the fears that I always have around these big Companies is that there’s always there’s always going to be big-company bureaucracy in politics and who takes credit for what and you know that sort of thing. And then I was in the when I was in the US over the summer, I had conversations with a couple of people who described working, you know, we were talking about Google and DeepMind at the time because it was during the week when everybody seemed to leave. And one of the ways it was described to me was you know, if you want if you think you’re going to change the world and you want to make AGI happen and so on so forth, then the most convex place where you can go and do that is at the frontier labs. And you know, doing the same thing at a big internet company feels like you’re designing a better toaster. which I mean it’s it’s not I don’t know if it’s a completely fair comparison, but yeah, and financially at the individual level, if you’ve done a super difficult PhD in something, you come out and you want to monetize that, then You know, today the most convex way to monetize that is probably to join Kimi ahead of their IPO, right? So, you know, there are those types of incentives at play as well. So, yeah, you know, so you know I do ultimately think the labs will be up there in terms of their ability to deliver these reasoning capabilities. and then the first and third party harness question ends up being kind of something that iterates in real time.
Grace Shao (56:58)
Right, right. And I think it’s kinda like going back to your earlier comment, like I think it’s what happened with Tencent when DeepSeek came out, they’re like, you know what, our Hunyuan is kind of meh. So maybe we’ll just focus on building products that are like a harness around it, like, but it just didn’t work because their own models weren’t good enough and then you can’t always rely on other people’s models, right? So these big tech are still going ahead with their own models now.
Robin (57:19)
Yeah. Well, they now own twenty percent of DeepSeek. So, you know, I think Tencent has the deep pockets and has the kind of strategic patience to be able to go down multiple routes where it, you know, you are doing your own model, you know, Hunyuan 3 was good and now Hunyuan 4 is coming soon. and meanwhile you’re still serving DeepSeek within your Yuanbao and your
Grace Shao (57:41)
Mm.
Robin (57:42)
WorkBuddy apps and so on, and over time And they use, for example, GLM in their ima app. and so yeah, they’ve always kind of done both. and you know, I guess it is reasonable that WorkBuddy allows them to see the reasoning traces of different models and that then somehow feeds back into their own model development, which is so yeah, I think Tencent’s in a slightly different position than a lot of the other guys in this conversation.
Grace Shao (58:11)
But wouldn’t Alibaba and ByteDance have the same kind of deep pockets?
Robin (58:16)
They would. But do they have the do they have the kind of social infrastructure and
Grace Shao (58:22)
Hm.
Robin (58:23)
The ability to pull everybody into a you know, an open third party harness? Whereas you know, if you look at Qwen Work, if you look at TRAE, these tend to be a lot more first-party-heavy. I don’t know if they will be as open in the fullness of time as Tencent. So yeah, they they y you’ve got these companies. And then, you know, each of these companies will then have their own internal kind of puts and takes in terms of who gets what. And so Yeah. Tencent historically
Grace Shao (58:49)
Yeah. They’re definitely all trying to follow the
Robin (58:51)
Had a bigger had a better track record of being kind of open and being the
Grace Shao (58:56)
Yeah.
Robin (58:57)
Somebody called them the benevolent gatekeeper of China Internet, which is yeah, it’s not a bad way to describe them, I guess.
Grace Shao (59:04)
Yeah, I think the other two are trying to follow the WorkBuddy route as well. They’re They’re all revamping DingTalk and Lark right now. So I think it’s like putting it into Qwen Work or something, and then TRAE—
Robin (59:13)
Yeah yeah yeah.
Grace Shao (59:14)
—TRAE and Coze were put into Doubao, and now it’s called Doubao Work, with Lark kind of grouped into it. Anyway, I wanna ask a question on recursive self-improvement. So I’m quite curious about what you think about the gaps between a lot of the labs. I don’t even want to position China versus the US, but if you have to put it that way, you know. A lot of times people are saying, you know,
Robin (59:35)
Yeah.
Grace Shao (59:36)
The Chinese advantage — sorry, the US advantage right now is that the labs are putting a lot more money into R&D and into figuring out how to go forward, right? And then the Chinese models some some ways are basically following their footprints and figuring doing the
Robin (59:48)
Yeah
Grace Shao (59:51)
Knowing the answer key to the homework. It’s that kind of analogy people are saying. But if there really is RSI, then would that change things? Then would these labs — sorry, the models themselves — just start figuring out How to improve themselves quicker and quicker and that gap would just shrink and compound with time or how do you see that?
Robin (1:00:10)
Yeah, I mean the I guess it depends on how you define RSI, but I guess the way I think about it is that there will be a gradient of different levels of how of you know automation and how to what extent AI can train AI and you know you can have them write kernels or whatever. But you know, going forward, can you ingest data and going back to the RL kind of example that I mentioned earlier, can you have AI write The verifiers and the gates and the success fail kind of conditions and you know have AI set up RL environments rather than somebody with a PhD doing it. so I think that will happen relatively quickly and then you know you start automating that process more and more. But then you still ultimately I think I’m very big on this. Like I still think you have taste and judgment be very important in terms of what kinds of verifiers you get the AI to to develop, and how you know there’s still ways to you know to do it better than the next guy rather than just have AI brute force everything. there will be some elements of that. But yeah I think you know and more and more it’ll be kind of you know the human supervising the AI doing more and more of it and but still kind of leaning on it and offering a bit of a steer in terms of where it where you want it to go and the types of behavior you want it you want to reward and so on. So yeah, I think I think there will be a gradient of how much AI versus human involvement you have in some of the model training. On the compute gap, then yes, if you have if, say, tomorrow OpenAI figures out RSI at a pretty high level and they’re able to iterate quickly, then that probably means that the gap between US and China widens again to some degree. You know, We seem to be in this Quite circular debate about whether it’s three or six or nine months and you know, I think the reality is that there’s it kind of oscillates. The US comes up with something and then China closes the gap again quite quickly. but yeah I think given how quickly information is kind of you know the flow of information between different parts of AI has been so rapid I think it’s, you know, over time the gap then, you know, maybe doesn’t widen infinitely.
Grace Shao (1:02:35)
And do you think this whole three-, six-, nine-month thing, end of the day, when we take a step back, surely it’s just so minor, no? Like how do we understand that?
Robin (1:02:44)
I think it matters to a bunch of people above our pay grade. You know,
Grace Shao (1:02:49)
Ha ha.
Robin (1:02:49)
A lot of it is being kind of hijacked in the media, in kind of geopolitical discussions and, you know, things of that nature. and if you are a frontier researcher, if you’re doing, you know, I think Anthropic has started to talk about medicine or drug discovery as a as one field that they want to be good at and If you’re trying to discover new drugs or trying to cure cancer, you know, which maybe we are now starting to do, then you do want, you know, the super frontier latest three months of model capability because and y you’ve seen enough c you know, like early-stage biotech whether these things either go up or down a hundred percent or w you know, whatever. But Because either a drug works or it doesn’t. And so for those types of use cases, yes, you do want the absolute frontier. If I’m just having a conversation with my Hermes bot, like if my model is three months out of date, does it really matter? I like to think it does. In reality it probably doesn’t.
Grace Shao (1:03:55)
No, it’s true. do you think there’s anything else that we need to talk about today, just for our audience to better understand the China AI landscape right now, where it’s at, or how to evaluate it?
Robin (1:04:07)
Yeah. Yeah, I’ll talk about something that’s a little bit of a tangent. But you know, one of these hills I’ve decided to die on is gaming. You know, I cover a lot of gaming both in China and Japan. There was this kind of moment where Google came out with Project Genie. I mean this was in this was in earlier on in the year where you know US software seemed to go down 5% a day every day. And Gaming got thrown in with that. And I think since then I think folks have come back a little bit and you know, back to the judgment and taste kind of thing. yeah, I continue to be of the view that gaming is actually remarkably difficult to disrupt and actually probably benefits from AI evolution over time. that’s probably you know and recently, you know, Google actually Was showing off one of my companies, Capcom, to demonstrate, look at how Capcom’s using AI and therefore, you know, AIs useful in the real world. And so yeah, I think I think the logic has kind of been turned on its head a little bit, but that’s that yeah, that continues to be the hill that I’ll die on when it comes to AI.
Grace Shao (1:05:21)
You think like gaming, filmmaking, these are all spaces where production cost is a lot lower, but you know, consumption will still be there. Is that kind of the thinking behind that?
Robin (1:05:32)
I think yeah, I mean production costs should come down as you automate more and more of these things. you know, you are starting to see AI video start to take over, you know depending on how complicated like you’re you’re not gonna have Nolan-level movies with generative AI anytime soon, but you are seeing short form video platforms essentially become AI-centric. so yeah, I think I think, you know, That will continue to grow, albeit the issue I’ve always had with multimodal models or video models is that, you know, the ceiling for commoditization, or the ceiling for at least you know, me not being able to tell the difference between one and the other is quite low. And so, you know, how do you differentiate video models from each other when that happens is kind of the thing that I’ve not been able to fully resolve. But yeah, if you’re just a maker of visual content, then this is great for you, right? You’re able to kind of Make stuff much more quickly. gaming in my mind is a lot more complicated because it’s not just about, you know, rendering of 3D environments or rendering of characters or whatnot. You have to have a storyline, you have to have, you know, action, music, fe you know, combat and so on and so forth that makes it more difficult. But yeah, or you know AI solves for production, you get a lot more output. I think the question in media and games and movies is like do people then care? Right? but then we’ve also just seen Niu Lai go viral. So maybe, you know, you need a more nuanced version of what people care about.
Grace Shao (1:07:10)
I like how we’re ending this conversation full circle. We started the conversation with having you look at the Ox Alpha thing and then the meme going around is the Niu Lai picture.
Robin (1:07:19)
Yeah.
Grace Shao (1:07:20)
So let’s see where that goes. I was like, is this a representation of niuma (牛马)? ‘Cause the end of the day we’re just all like worker bees and in Chinese we’re all horses and cows. But he said probably not. That’s not where the meme comes from.
Robin (1:07:35)
I’m gonna leave that alone. I’m gonna leave that alone.
Grace Shao (1:07:43)
How do you use AI in your own research, but I want to ask you like there’s just a lot of AI tracking tools out there. There’s a lot of scattered data. AI itself is obviously very broad to even kind of you know, it’s just to say what are you tracking? So actually the question for you is like, what are you using for your own evaluations or what do you track? Like, are you looking at OpenRouter, ModelScope, GitHub? Like what, what do you use to understand? Where demand is going, how compute is used, which models are good, etcetera.
Robin (1:08:15)
Yeah, I mean we track all the obvious things. I mean I’m not sure that’s differentiable necessarily, but you know, OpenRouter, Hugging Face, GitHub, ModelScope. We’ve set up kind of bot routines to try and scrape these things every so often. And, you know, we have these gigantic Excel files of, you know, daily data, weekly data. And then that at a high level paints a picture of you know who’s actually engaging with some of these things, who’s actually, you know creating repos, who’s then downloading the models, who’s doing this, that, and the other. one of my pet peeves with OpenRouter is that, you know, it’s a really, really small chunk of the market. a lot of people, especially on in the investment world, are kind of overindex on it as a as an indicator of and you see the media do it and say things like, You know, Chinese models are now two-thirds of consumption or something. It’s not, it’s two-thirds of consumption on a relatively small part of the market. but fine. You know, we do track it and we do look at kind of, you know, which apps are being used, be it kind of Z Code or some of the other stuff. So we do all that. I guess one thing that we do, which I don’t know if anyone or many other people do, is we do run our own evals. So we When a new model comes out, one thing that we do is we run it through or my bot runs it through these benchmarking tests and you make them solve tasks and it gives me I mean there’s a there’s a variety of reasons to do this, but effectively gives you a more first-hand view on whether you think something is good and then I tend to try and use them day to day because that gives you then potentially a more varied kind of read on Model feel and usefulness beyond just making it crack hardcore software engineering tasks. which, you know
Grace Shao (1:10:13)
It’s not for everyone.
Robin (1:10:14)
Well, that, but also it’s you know, these labs presumably then train on something like Terminal-Bench or SWE-bench. No one’s gonna train on my own day to day nonsense. So yeah.
Grace Shao (1:10:26)
Do you have any other tips for how us like research-driven people or our jobs are just here typing away on our computer how we can better use AI in our research Process?
Robin (1:10:40)
I will say it’s an iteration process. Like I think even you know, you have to kind of use it and be hands on, figure out, you know, what is useful for you the most and and you know find ways to kind of make it multiplicative, right? And just not just kind of have it summarize the news, but also, you know, we’ve built different Constructs on top of my data extraction funnel to try and you know have it actually analyze what’s going on, give it, you know, send me kind of reads on different things that I care about at any given point. And then I use it to iterate some of the stuff that I do day to day.
Grace Shao (1:11:24)
It can be as smart as how you make it to like how smart your inputs are, really. Yeah.
Robin (1:11:29)
I think so. And you know, what I found is you go around in circles, like you give it some stuff, you kinda work out whether it actually can do that or not. And sometimes you know, the answer’s not always yes. and yeah, and you know, try and kind of extend or you know, or try and give it you know more and more things to do until you so you know, just organically there will be use cases that kinda pop out from every so often based on my experience.
Grace Shao (1:11:59)
I don’t condone the fact that you weren’t spending time with your daughter full on one on one and actually on your voice AI tool. But my husband’s gotten to a point where he’s wearing his Apple Vision Pro and he has like seven of his agents in a line and he’s just sitting there like controlling them. I’m just like it’s gone too far. Like I think if there’s gonna be like a pushback by all the kids and wives at this point, not to overgeneralize. Now, look, what has really changed for you or your view on over the last six months? you know, has something fundamentally shifted in your research or something your understanding of AI and the technology.
Robin (1:12:39)
My understanding of AI seems to change every two weeks. So, you know, that the whole process of being in January, staring at these two IPOs, and then you know, fully going down the rabbit hole of trying to understand them and keeping you know, trying to keep track of everything, build a network around, yeah, I mean it’s it’s been it’s been wild, how much everything has changed. I’d love to be able to distill it down to one thing, but I think the reality is just, you know, everything has changed and the thing that I’ve discovered actually that’s been interesting is you know, I run a small team and there are people who normally work for me and in the old world I used to kind of staff them on a small handful of things a day. And then they will go away and do it and they come back and I would have to make, I don’t know, three or five consequential decisions a day or, you know Have kind of deep, deep conversations with myself about how to do something. Now, increasingly, with AI — I mean they’re still doing that, but at the same time I’m knocking around ideas in my head and I’m sometimes using AI bots to help me kind of process what I think. and they come back so quickly that I find myself having to make five consequential actions an hour. And then at the end of the day I’m like, you know, my workload has increased as a result, which is I don’t know if that was the desired outcome, talking about AI making people’s lives better. But there you go.
Grace Shao (1:14:15)
I think it’s just because you’re too much of a doer and too competitive. I think for people who are high-agency people, they are doing more and they’re experiencing AI fatigue. For people who wanna just clock in, clock out and just do the three things they were told to do, their lives are made easier. You know, it’s like you know, writing three emails used to take maybe like a day, now it takes twenty minutes or less.
Robin (1:14:35)
Then I wouldn’t be on this pod, so there are upsides.
Grace Shao (1:14:39)
Look, one last question for you. I asked everyone, what is one differentiated view you hold or you know, something a bit non-consensus you think?
Robin (1:14:49)
Yeah, I think the gaming thing is probably the most controversial. you know, or the maybe the broader idea of I think if you want I think it’s it’s the whole idea of of judgment and taste where, you know, there’s still this belief and it’s almost a religious belief at this point, which is because I don’t know how to disprove it either. which is I think on some level it’s you know That will always be valuable, the idea that I’m still of the view that it’s quite difficult to get AI models to spit out something that’s outside of its own distribution and training data. And so, you know, on some level you still have human input being the arbiter of something that’s ninety percent of the way there and something that’s truly kind of differentiated. So yeah, I vote at the end for humankind.
Grace Shao (1:15:47)
But then I have a follow-up question on that. It’s just like I think it’s easy for people who’ve built up taste or, you know, like yourself, you’ve been in the industry for long enough to build up your own taste, your own judgment. How do people without that kind of experience build up human taste still when they’re joining the workforce or they’re growing up like our children are growing up at an age where, you know, AI is natively embedded in everything they do? So, how do you still differentiate that taste? Because You know, now like we’re seeing these like parallel structures of sentences everywhere you go. It’s driving me crazy. Even Instagram ads are like, you know, they have like these very obvious parallel structures and it’s like driving me crazy. I’m like, dude, like this marketing associate did not write this, but I think what you’re
Robin (1:16:29)
Okay.
Grace Shao (1:16:29)
Seeing is people without that taste judgment and now just copy and pasting whatever AI spits out.
Robin (1:16:36)
Yeah. I mean we’re gonna end this conversation like every good Asian parent and talking about parenting at the end of it. no, look, I mean it is a conversation I have with myself. Like how do you educate somebody or how you know, whether it’s your own kid or whether it’s somebody that you work with or whatever, then, you know, how do you develop taste from first principles and Yeah, I don’t know. I don’t have a great answer, but I do think exposing yourself to original content and doing things the hard way, at least in the beginning, is still important.
Grace Shao (1:17:16)
Well, thank you so much for your time. Very generous with your time today, Robin. I really appreciated your insights and everything.
Robin (1:17:21)
Appreciate our conversation.
AI Proem is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
Get full access to AI Proem at aiproem.substack.com/subscribe
Avsnitt sparat!
Du hittar sparade avsnitt på Mina sidor.
Kunde inte spara avsnitt
Något gick fel. Försök igen.