AI Proem Podcast
Avsnitt

From 3D Design Software to Spatial Intelligence: Manycore’s Next Chapter

Dela

In this episode, I speak with Bei Shen, CFO of Manycore Tech, the Hangzhou-based company behind Kujiale in China and Coohom overseas. Manycore started as a cloud-based 3D design software company serving designers, furniture brands, retailers and property developers. Today, the company is expanding into spatial intelligence, using its 3D data, simulation capabilities and software to explore applications beyond design.

We talk about Manycore’s evolution from startup to public company, including its IPO earlier this year and how the company’s story has changed since going public. While its core SaaS business still accounts for the majority of revenue, Manycore is increasingly positioning its proprietary 3D data and technology as a foundation for spatial intelligence and new applications.

We also dive into SpatialVerse, AholoWorld and Manycore’s work in robotics and embodied AI. Bei explains how the company thinks about spatial intelligence—not simply as a data business, but in terms of the systems and simulation environments needed to help AI understand physical space. We discuss potential applications in robotics, game design and filmmaking, as well as the question of how much intelligence different types of robots actually need.

Finally, we discuss Manycore’s global expansion, partnerships and long-term strategy. We explore whether spatial intelligence will become a market dominated by a few global platforms or remain fragmented across industries and geographies, and what Manycore sees as its role as more robotics companies begin building their own physical AI systems.

The AI Proem Podcast is under the AI Proem newsletter, which has over 12k followers globally. To learn more about China AI, the business of AI, and how AI is impacting businesses, please check out the newsletter here and more insightful conversations here.

Chapters

01:18 Leaving Investment Banking for a Startup in Hangzhou

05:29 From Silicon Valley to Going Public in Hong Kong

06:19 First of the Six Tigers to IPO

07:28 3D Shift Toward Spatial Intelligence

14:20 Data, Simulation or System: What Is the Product?

18:14 Learning More of SpatialVerse and AholoWorld

31:59 The Role of Spatial Intelligence in Robotics

34:04 How Manycore Fits Into the Robotics Stack

38:11 Global Expansion and Strategic Partnerships

42:31 Will Spatial Intelligence Consolidate or Fragment?

46:20 Manycore’s Focus and Strategic Pillars

AI-generated transcript (for reference only)

Grace Shao: Hi, Bei. Thank you so much for joining us today.

Bei: Hi, Grace. Good to be here.

Grace Shao: Yeah. So, you guys are one of the hottest AI companies that listed in Hong Kong this year, and a lot of people have a lot of questions. But to start with, tell us about yourself. When I was learning about your background, I thought it was quite fascinating. You’re almost like a Joe Tsai story. You had a very successful finance background and a successful career in Hong Kong, and then you decided to jump over to Hangzhou to join what was still a relatively unknown startup. Tell us what made you want to make that jump and leave your cushy banking role. What was the spark about this company for you? And tell us a little bit about where the company is right now.

Bei: Sure. My name is Bei Shen. I’m CFO of ManyCore. I joined the company in 2019. Before that, I was an investment banker for 14 years. I worked for Citigroup in New York, then moved to JPMorgan in Hong Kong. The last nine years of my banking career were at Goldman Sachs.

Around 2018 or 2019, I started thinking about what I wanted to do with the rest of my career. Traditionally, I focused a lot on clients in more traditional industries. I covered companies in the power, mining and energy spaces. It’s an interesting job, but the sectors are relatively traditional.

So I started asking myself how I could get more exposure to technology and internet companies. For me, it was very difficult to switch industries within the bank, so I started looking around for opportunities. Luckily, ManyCore was looking for a CFO. After talking with the founders and the team, I found the company very exciting, so I joined in 2019. I can’t believe it, but it’s been almost seven years now.

Grace Shao: Yeah. And I know you recently took the company public, but before we get into all that, tell us about your three co-founders, because they have quite interesting backgrounds. They’re quite young. They came back from Silicon Valley, bright-eyed and wanting to start something in China. Tell us about the vision they had in the early days and where it has led you now, roughly 15 years later.

Bei: Yes. The company was founded around 2012. The three founders were classmates at UIUC, which has a very strong computer science program in the U.S. Our chairman, Victor, and our CEO, Chen, actually went to the same undergraduate university, Zhejiang University, which is also where our company is based. We still recruit a lot of people from Zhejiang University, which is a great school. And our CTO went to Tsinghua.

All three of them studied at UIUC in fields related to computer vision and high-performance parallel computing. After graduation, they all went off to cut their teeth in Silicon Valley. Victor worked for NVIDIA for a couple of years, Chen worked for Microsoft, and our CTO worked for Amazon.

They were all working in Silicon Valley, but they wanted to come back to China and participate in this exciting market. Back then, Victor had this idea of putting GPUs on the cloud to serve more clients. He was working at NVIDIA on the CUDA team, so he was involved in the early days of figuring out how to put compute on the cloud and serve more customers. Obviously, this was before AI became what it is today.

They built a demo and came back to China. Luckily, the Hangzhou government was welcoming overseas graduates and helped them start the company.

The original idea was very simple: they wanted to put GPUs and compute in the cloud and make that compute available to more people. In the beginning, it was very difficult because AI hadn’t taken off yet, autonomous driving wasn’t in full swing, and crypto wasn’t either.

Luckily, they found a very interesting application in interior decoration. In the old days, if you used on-premise software, it could take a very long time to render a photorealistic picture. With their technology, you could put that computation on the cloud and use multiple GPUs to accelerate the rendering process. That enabled users to create photorealistic renderings in minutes. Now it’s seconds.

That really changed the industry in a big way. So that’s how they started. It was fundamentally a technology company trying to find applications for its technology.

Grace Shao: How would you describe the company today? How would you position ManyCore in two or three sentences? Clearly, it’s no longer just about Kujiale and 3D interior design.

Bei: Obviously. The company has had 14 or 15 years of history. Before 2023, we were basically the largest 3D design software provider for interior design. But since 2023, the company has increasingly focused on spatial intelligence.

To put it very simply, we’re trying to help AI perceive, create and eventually act in a three-dimensional world, so that AI can eventually move from the digital world into the physical world. That’s where the company is focusing right now.

Grace Shao: Perfect. I think you were definitely one of the hot IPOs earlier this year. You went public in April and were one of the first of the Hangzhou “Six Little Dragons” to list. It felt like a point of pride for Hangzhou and for this new wave of Chinese AI companies. What did going public mean for you and for the company?

Bei: Obviously, it’s a big milestone for the company. We raised fresh capital to fund our future growth, especially in spatial intelligence. We need more compute and we need to hire more talent.

But from a business perspective, it also put us on the international radar. We already have many international clients, but it can still be difficult for a Chinese technology company to sell products to overseas customers. Being a public company, with your company story and financials becoming more transparent, definitely helps a great deal in promoting ourselves and selling our products in markets outside China.

Grace Shao: I want to double-click on something you said earlier about how the company evolved. When you filed the prospectus, I went through it, and it was still mostly focused on your 3D interior design technology. Now you’re clearly pushing a new narrative around spatial intelligence, which frankly wasn’t emphasized nearly as much even a year ago when you filed the prospectus. Things are moving so fast.

Tell us about how that shifted and why you had this moment of pivot. Was there an epiphany during the process of going public, or after you went public? Did something hit you where you realized there was this gold mine you were sitting on? Tell us the story behind that.

Bei: Sure. That’s an interesting question. Just to go back a little bit in terms of our IPO history, we really started preparing for a Hong Kong IPO in the third quarter of 2024. Then we filed our prospectus on February 14, 2025.

The IPO process is relatively lengthy for Chinese companies because every company going public needs approval from Chinese regulators. For us, it took a bit longer because of our structure. We finished the IPO in April this year. So looking back, the process took almost a year and a half.

Obviously, both the company and the industry changed enormously between the day we started the IPO process and the day we actually listed.

Our thinking was that it would be unreasonable, or even impractical, to keep updating the prospectus every time the company changed because this industry moves so quickly. So we made a decision to keep the discussion of the new business and products relatively minimal.

That’s why, when you read our prospectus, you see a lot of disclosure about our older, existing business, which is obviously still important. But the new businesses were changing so much that we didn’t go into as much detail.

After the IPO, we started talking to more analysts and investors and trying to give them a more updated picture of where we stand in spatial intelligence.

Grace Shao: For sure. So what are the top-of-mind questions or areas of interest you’re getting from investors right now about the business?

Bei: This is a very frontier area. Large language models have obviously received a lot of attention over the last couple of years since ChatGPT came into existence. There have been many advances in model capabilities, coding, image generation and video generation.

But we’re focusing on a relatively different type of AI. We sometimes call it physical AI. As I said, we’re trying to help AI understand the 3D world, which is very different from reading text and giving you an answer, or generating a picture or a video.

One of the challenges in our space is that we don’t have nearly as much data as large language models do. They can access internet text and enormous amounts of video. In our space, the amount of data is several orders of magnitude lower and much less dense compared with text, pictures or video.

That’s why it’s difficult. It’s very hard. People also need to spend some time understanding what we’re doing because I believe we’re working at the frontier of what could be the next wave of breakthroughs in AI.

Grace Shao: My understanding, according to your public filings, is that roughly 90% of your revenue is still coming from the traditional business, particularly Kujiale, and that’s really funding the new initiatives.

As you move into spatial intelligence and physical AI, you touched on data as the bottleneck. But data is also part of your moat, right? You’ve had more than a decade of experience working with 3D data. Tell us more about the connection between your traditional business and this new business.

Bei: Sure. I wouldn’t say data is the only reason we chose spatial intelligence, although it’s obviously a very important aspect of the business.

There are many connections between our traditional Kujiale business, or Coohom internationally, and what we’re focusing on today.

Even 10 or 12 years ago, we adopted a very integrated technology architecture. We bought our own GPUs and did rendering using our own GPUs in order to provide the service to customers worldwide.

That’s actually quite similar to what large language models or 3D models are doing today. You’re utilizing compute to provide products and services to people around the world over the internet.

So that’s one connection in terms of the technology lineage. We did a lot of hardware-software optimization to make sure rendering could be provided at the lowest possible cost. Similarly, if you want to do inference today, even if you have a very good model, you still need to keep inference costs low in order to remain competitive. That’s something we’re obviously very good at.

Data is also very important. We’ve accumulated a large amount of data. In hindsight, it’s fortunate that, compared with the on-premise software that came before us, we had all the data on our cloud platform. That wasn’t necessarily by design at the beginning.

But in today’s world, 3D data is extremely valuable and very difficult to obtain. Because of our cloud-based architecture, we’ve been able to accumulate a large amount of data, especially structured 3D data, which is critical for training 3D understanding and 3D models.

So between our technology lineage and our data library, I think we’re in a very unique position to explore spatial intelligence.

Grace Shao: So you have a lot of 3D data, but this is spatial data. It’s not necessarily the movement or motion data people talk about needing for robotics training.

But when I spoke to your team while visiting Hangzhou, it sounded like a lot of your clients may actually be robotics companies. What are you providing to them today? Is it a 3D intelligence system? Is it data that helps robots operate better in physical space? Or is it a bit of both?

Bei: This has an interesting history. I think it was back in 2022 or 2023, during the pandemic, when nobody could really go anywhere. We were all stuck in offices or at home and couldn’t travel abroad.

We received an email from Silicon Valley from one of the large technology companies. They came knocking on the door and said, “I heard you guys have some interior-setting data.”

We said yes.

They said, “We’d like to buy some.”

At the beginning, we thought it was spam or some kind of trick. But it turned out to be a real client doing research.

We didn’t think too much about it. We struck a deal and helped put together some synthetic data they required. We made some money, not much.

Then the next year, another technology company came and asked for something similar. This time, we took notice. We thought, okay, there must be something valuable in our data.

So we started asking these U.S. clients, “What are you actually doing with our data?” We’d had it for over a decade and hadn’t really thought too much about it. Luckily, we hadn’t deleted it just to save storage costs.

Only then did we find out that these were large technology companies in the U.S. training robots. They needed synthetic settings in which they could test and train their robotics policies.

That’s when we realized we were sitting on something interesting and valuable. We started thinking about how we could better commercialize the data we had.

Obviously, we’re not satisfied with simply providing raw synthetic data. Right now, we’re speaking with customers both in China and the U.S. and trying to help them train and evaluate their policies more effectively.

Eventually, we also hope to train our own model.

We believe that if you want a world in which robots, or physical agents more broadly, can become fully autonomous, they need their own brain. The ability to perceive and understand physical surroundings, reason about them and act within them may require spatial intelligence. That’s something we’ve also started working on ourselves.

So right now, we’re still providing synthetic data to customers. We’re also trying to train our own model, which we eventually hope can be put into intelligent robots or embodiments of different forms so that they can really act in the physical world.

Grace Shao: That’s really interesting. There’s a bit of serendipity there. Things just happened and you were there at the right time with the right data.

I was reading through your materials. There’s something called SpatialVerse, and you’re also releasing HoloWorld. What are these things for? Tell us more about these products.

Bei: Sure. These names keep popping up. Sometimes I get confused as well because things change so fast.

As I said, in 2022 or 2023, we started selling synthetic data solutions to robotics companies. Later, AR and VR companies also came to us asking for something similar because they need data to train goggles or glasses.

We put this type of business together under the name SpatialVerse. That’s one line of activity and business we’re developing.

As I said, eventually we’d like to train our own models and put them into robots.

Another branch of research we’re working on is helping agents create synthetic worlds. This is somewhat similar to what Dr. Fei-Fei Li’s company, World Labs, is doing.

Basically, using a simple prompt, text or pictures, you can quickly create a virtual space where you have geometric information as well as information about the objects within that 3D space.

People talk about “world models” a lot these days, and sometimes the term is misused or misinterpreted. But the way we understand it, in order to have this capability, you really need a model that can generate a 3D world.

It’s not just a continuation of pictures. There are very good models today that can give you 20 or 30 seconds of short video. But what we’re after is the ability to use computing technology to generate a 3D world where the objects within that world have certain physical properties, and where you also understand the geometric relationships between those objects.

That’s critical for robots eventually being trained inside that world.

Grace Shao: So rather than a video-generation model, which is where a lot of multimodality efforts are going right now, you’re really trying to create 3D spaces. It almost feels like creating a little Sims world.

What are the use cases then? Off the top of my head, there might be game design, filmmaking and, of course, robotics training. What do you think people are missing when they think about what this tool or product could eventually be used for?

Bei: That’s a great question. Again, this is relatively new.

Historically, 3D design has been a relatively niche market. Before us, you had all these on-premise 3D design software products, but they’re difficult to master and use.

Compared with picture editing, 3D design has historically been quite niche. It was mostly used in architecture, industrial product design and VFX.

What we did was make the 3D world easier for ordinary people to generate.

One area where we’re already helping customers is media. In China, one-minute and two-minute micro-dramas have become very popular. You see many of these productions on Douyin, and a lot of the content is already being generated by AI.

We’re helping some of these creators produce spatially consistent 3D worlds. If you rely purely on video generation today, you can get hallucinations after a certain amount of time. Objects start floating around. You leave a room, come back, and suddenly the objects are missing.

Using our product, these producers or content creators can maintain a spatially consistent 3D world and produce something higher quality. You don’t have the same spatial hallucinations you get from video models.

The other application is robotics training and evaluation, which we’ve already discussed.

It’s inconceivable that you can train robots in every possible physical environment. Thinking through all the corner cases would be too expensive and too time-consuming.

So in order to train and evaluate these policies, I believe it’s critical to have virtual worlds where you can put robots through testing virtually.

Those are two areas where we’re already putting our technology to use. Eventually, I think there will be many more applications.

Grace Shao: That’s really interesting. I want to double-click on the micro-drama example.

To help me understand, are people using your technology in parallel with something more traditional like Seedance? One is more for aesthetics and one is for spatial control? Is it layered?

Or could your technology eventually compete with and replace a more general-purpose video model like Seedance?

Bei: That’s a great question. Right now, the way we serve customers is really a layered approach.

We already have a product out that you can try called LuxReal. We rolled it out about a month or two ago.

Basically, you can upload a script, and it’s an agent that helps you produce a 30-second, one-minute or two-minute micro-drama based on that script.

First, you use our product to create the 3D world. For example, if you want to shoot a micro-drama set inside an ancient Chinese palace, after reading your script, we’ll generate that palace for you. Let’s say you want two rooms inside Beijing’s Forbidden City. We can create those rooms, and then you can move your camera around within that space to shoot the scenes.

Our customers don’t only use our modeling capability. They also use Seedance because we don’t actually do the video-generation part right now.

We help you create the spatially consistent 3D setting, and then you put Seedance on top of that. You can very quickly produce a one- or two-minute micro-drama.

If you only use Seedance, you may have to do a lot of editing afterwards because certain things don’t look right. You have to spend manpower editing out hallucinations.

With our product, the room is always there and the table is always there. The table isn’t going to change when your camera changes.

So it’s basically a very efficient tool for these micro-drama producers.

Grace Shao: That’s very interesting. So essentially, you’re producing one asset layer in the broader workflow for these creators.

It’s funny because when I was talking to people at Kling and Kuaishou, they said people don’t necessarily care about inconsistencies yet. Sometimes the cat turns out white, sometimes the cat turns out black. There’s definitely an understanding that AI content, as of now, isn’t that sophisticated.

I want to shift the conversation to models and robots.

I’m going to put you on the spot here. You mentioned Dr. Fei-Fei Li. Yann LeCun and Fei-Fei Li are both world-renowned scientists working on something around this realm. They’re both trying to push forward ideas around world models. Obviously, this is still in a very nascent stage.

How do you see the field? And frankly, how do you see your company’s position among all these global competitors, which have a lot of technological influence and obviously a lot of capital behind them as well?

Bei: Great question. I wouldn’t say we’re directly competing yet. As I said, this is a fast-evolving and changing industry, and everyone is still doing a lot of exploratory work.

Dr. LeCun’s approach, to be honest, I don’t understand in great technical detail because I’m not computer-science trained. I’ve read about it, and one apparent benefit of his approach is that you may not need as much compute to come up with these highly realistic representations.

We haven’t paid too much attention to his work yet, but obviously we’ll be very interested to see what he develops.

Dr. Fei-Fei Li’s approach is more similar to ours. Both companies are trying to figure out an efficient model for creating a 3D world where you not only see the world as we see it, with the correct textures, sizing, depth and perception, which is the rendering side and something we’re very good at, but where you also embed more information.

We’re trying to add physical properties such as friction coefficients and how wind behaves. As customer requirements evolve, we’ll try to put more and more information into that model.

Eventually, it will be interesting to see what applications emerge outside media and robotics training when people have these kinds of worlds available.

Grace Shao: I listened to one of your CEO’s interviews, and he said that what you’re trying to do is create spatial intelligence that can help translate physical space so LLMs can better understand it.

He also discussed how world models shouldn’t necessarily be produced by every individual robotics company, or at least shouldn’t be viewed as interchangeable, because each use case can be so different.

So I guess my question is: how should we think about a robot being used to remove blood clots, where the work is incredibly meticulous, versus a robot designed to lift heavy objects? What kind of 3D data and 3D model does each need?

It seems like “world intelligence” or “world models” is too broad a category to serve every single demand in the space right now.

Bei: I’m not 100% sure which episode or interview you’re referring to, but I think what he was probably talking about is the robotics industry today.

Obviously, the level of optimism varies depending on who you talk to. But based on our conversations with the industry, a general-purpose humanoid robot is still pretty far away.

I wouldn’t say it’s next year. Maybe it’s five years, maybe it’s 10. But just imagine a humanoid being able to do 100 chores in your household. I think that’s still quite a few years away.

There are so many challenges. One of them is exactly what we’re working on: can you teach a robot to understand different settings?

The minute it walks into a room, can it immediately understand, “Okay, this is a table, this is a desk, and these are the relationships between the objects”?

We’re still working on that.

Building a general-purpose, fully autonomous humanoid robot is hard.

But if you narrow the problem down to more vertical or specific-purpose robots, I think that’s more doable. The level of complexity and intelligence required is much easier to achieve at this stage.

I think the industry is more likely to evolve through more and more specific-purpose robots. One robot might pick up boxes. Another might help with laundry. I’m just giving examples.

That seems like a more likely roadmap than trying to immediately build one general-purpose robot with omnipresent capabilities.

As you build more and more of these scenario-specific robots, maybe eventually you arrive at a stage where a more general-purpose robot becomes possible.

That’s how we see the world and how we’re tailoring our R&D efforts. We’re not trying to go after one extremely general spatial-intelligence model right now. We’re trying to crack these silos one by one.

Grace Shao: So which verticals are you focused on today, out of all the different kinds of robots you’re serving?

Bei: To give you a few examples, we think household robots are probably difficult, at least in China, because Chinese households tend to live in relatively small spaces. The margin for error is extremely small.

So right now, we’re focusing on helping robots in more industrial settings.

For example, in warehouses, we can teach robots to quickly understand the warehouse because the level of complexity is relatively lower than in a family setting.

We’re also helping some robotic dogs patrol power stations. They need to walk around, identify anomalies, record them and report them.

We believe these are some of the low-hanging-fruit applications today, where we can teach robots or robotic dogs to perceive and understand the 3D world.

Grace Shao: Do the economics make sense right now? Frankly, if you’re trying to replace relatively simple tasks or labor, especially in China or elsewhere in Asia where labor costs are relatively low, does it make economic sense?

Bei: That’s a great question. The economic equation is definitely important as we put more effort into this.

Some of the key areas we’re trying to explore are places where it’s dangerous or costly for human beings to operate.

For example, around high-voltage power stations or transmission lines, it’s definitely safer to have robots patrol instead of human beings.

Or in remote areas and underground mines, if you have water leakage or some geological situation, it’s safer to send robots and perhaps drones to inspect first rather than sending a human rescue team directly.

Grace Shao: That makes sense.

But my understanding is that you’re purely on the software side right now. Does it make sense for these humanoid robotics companies to pay for your service and technology, or does it make more sense for them to train and build their own models internally?

How do you view that? There are obviously different camps. Some people say OEMs can build different types of hardware while companies like yours provide the intelligence underneath. How do you see that trend?

Bei: Right now, we’re only focusing on software, as you correctly pointed out.

We’re trying to be model-agnostic, and we’re also trying to be embodiment-agnostic.

Basically, we’re trying to develop technology that different robotics companies can use to train and evaluate their policies.

Obviously, we’re not there yet, but that’s our goal.

We’re not trying to build robots or robotic dogs ourselves. We’re trying to help these companies find a very cost-effective way to evaluate their policies. That’s our approach right now.

Grace Shao: What do you think people are getting wrong or misunderstanding about the industry today?

In spatial intelligence and robotics, there’s obviously a lot of buzz and a lot of hype. Unitree has been getting a lot of attention as well.

Do you think there’s too much hype right now and that we should be more cautious because progress is still going to be slower than the public expects?

Or do you think the misunderstanding goes the other way, and people are underestimating how quickly this technology could proliferate in niche use cases and eventually reach consumers?

It’s a big, open-ended question.

Bei: Sure. From my perspective, I obviously believe this technology has very broad applications in the future. But the capabilities have to get there first, and the cost has to be low enough for the technology to proliferate.

I believe spatial intelligence is a critical part of human intelligence. It’s almost innate to us. Dr. Fei-Fei Li has made a very good argument around that.

After thousands of years of evolution, human beings can see things and quickly understand their geometric and 3D relationships.

That’s something large language models don’t really have today, but it’s critical if AI is going to operate in a physical context.

Right now, it’s great that AI can solve math problems or write poems. But can you actually ask a robot to do your laundry? Can you trust it to do all these tasks?

Right now, we’re not there yet. But I believe we’re on the way.

I wouldn’t necessarily call it a misunderstanding. I think the difference in opinion is really about how long it’s going to take.

As I said, one of the biggest challenges facing our industry is data. We need to find smart and cost-efficient ways to obtain more data because models are a product of that data.

Large language models are really the product of compute multiplied by data. Our industry is no different.

That’s why we’re thinking about different ways of capturing more 3D data. We’re working with different hardware companies. Robotics companies are one category. Scanning companies are another.

We hope to have more hardware companies work with us so that more users can use our technology to capture or generate 3D data.

In the long run, we need a flywheel where applications, data and models all improve in tandem, level by level.

But right now, the flywheel isn’t flying yet. We’re working very hard to push it forward.

Grace Shao: That makes a lot of sense. More users mean more use cases and more scenarios, which give you better data. Better data improves the models, which then lets you serve clients better.

On partners and clients, you mentioned earlier that you already have quite a global footprint. I think that’s fairly unique among Chinese companies that are trying to go global today.

When you think about international expansion and distribution, what’s most important? What kinds of partnerships are you looking for?

You mentioned that some large Silicon Valley technology companies are already clients. How should we understand those relationships?

And more importantly, do you face localization as a bottleneck, or is that less of an issue in your particular sector?

Bei: Great question.

Putting it in the context of Kujiale or Coohom, localization is actually very important.

For example, if you want to sell an interior-design product in the U.S., the industry is very different. China and the U.S. both consume furniture and decoration services, but the industry relationships and dynamics are very different.

Localization is therefore critical.

Just to give you a very simple example, even the measurement systems are different. China uses meters, while the U.S. uses feet and inches.

That’s a tiny example, but there are many localization changes you need to make in order to fully satisfy the local market.

What’s interesting is that this has changed quite a bit in the AI context.

ChatGPT basically became global overnight. I think one reason is that the model itself became so powerful.

You can almost think of the model itself as the product. You don’t necessarily need to build many layers of user interface on top of it.

We’re beginning to see that in our space as well.

If you have a very powerful world model, for example, you can generate 3D objects or scenes relatively easily. There may still be differences in language, but in terms of usability and application, it becomes much easier to promote globally.

That’s a big opportunity for us.

Eventually, we hope to offer a product that has a global appeal similar to ChatGPT, where you have users all around the world.

Obviously, that’s not easy. First of all, you need to come up with an extremely strong model that can actually serve clients and users worldwide.

Grace Shao: So what I’m hearing is that on the consumer-facing side, something like Kujiale obviously requires more localization.

But if you increasingly position yourself as a B2B support technology or an infrastructure layer underneath consumer-facing products, you may need less localization. Is that a fair understanding?

Bei: That’s a fair summary.

Coohom, by the way, is the international version of Kujiale. It’s not just a language translation. We’ve adapted Coohom depending on which market we’re entering, so we made a lot of localization changes.

But for a new product like LuxReal, the micro-drama product, we didn’t have to do much localization besides language.

That gives you an idea of how, in this era, products can become international much more easily because the underlying layer becomes extremely powerful and important, while the application layer on top can be relatively simple.

Some clients can even develop their own applications based on our technology.

That’s how we envisage the future.

We’d like to develop very powerful models and offer them through APIs or SDKs. People can then do their own development and build secondary or tertiary applications on top of the model.

That’s a change in paradigm compared with the past. As a software provider, you used to have to build many of these applications yourself.

Now users and customers can use things like vibe coding to build many applications themselves. We don’t necessarily have to do all of that anymore.

Grace Shao: That makes a lot of sense.

So just one last question on this part: should we understand the future of spatial intelligence as being more fragmented and vertical by sector and use case, rather than by geography, compared with how software evolved during the internet era?

Bei: I would tend to agree with that assessment.

As I said, because the data is so difficult to obtain, I think the industries where we can establish these small data flywheels will develop more quickly than others.

So I believe it’s going to be a more fragmented landscape compared with large language models, where eventually you may have fewer than half a dozen truly global companies. There are obviously more today, but I think LLMs will ultimately become quite concentrated.

In large language models, the data is relatively open to everyone because everyone has access to the internet.

A lot of the competitive landscape is therefore determined by who has more compute or who has the best talent and algorithms. Large companies have a huge advantage in that environment.

Our space is different.

There isn’t a universal 3D data library where everyone can simply start working on the same dataset.

First of all, obtaining the data itself is a challenge.

We obviously have an advantage because of the work we’ve accumulated over the years, but eventually we still need to find more and more methods of getting additional data.

So I think the game is somewhat different from large language models.

Grace Shao: I want to take a step back.

You guys are based in Hangzhou. Like you mentioned earlier, there was an effort in Hangzhou to attract people to come back because of how strong the ecosystem is.

Obviously Alibaba is there, Ant is there, there are a lot of e-commerce players, and many of the startups that came out of Hangzhou over the last decade have somehow been related to e-commerce.

It’s interesting that you didn’t get sucked into that orbit.

When I was reading about your story and looking at the earlier days, I thought it was funny because you could very naturally have gone into e-commerce staging and 3D content creation. That could have been a logical path for serving domestic clients, especially given that you were based in Hangzhou.

What was the thinking behind not going into that vertical?

Bei: We’re no exception. We tried e-commerce. It didn’t work out.

Grace Shao: I love the candidness.

Bei: We’re no exception.

Going back to Kujiale’s early days, we came up with this interesting software for designers. The natural next step was: can we sell furniture?

We tried. It didn’t work out.

I think that’s probably largely due to the genetics of the founders. They weren’t from that industry. They’re not e-commerce experts.

It’s also simply the nature of furniture and home decoration. It’s very difficult to commoditize. It requires a lot of service. It’s not like selling a book or laptop through e-commerce.

Even Alibaba tried, and I don’t think the results were very satisfactory.

So we dabbled in it. We burned some investors’ money, but not too much.

Then we realized, okay, it’s not for us.

We came back and said, we’re going to focus on software.

And lo and behold, we found this opportunity in spatial intelligence.

Grace Shao: Definitely. That makes a lot of sense. Furniture isn’t an easy thing to sell. It’s tailor-made, personal, huge and bulky, and logistics aren’t easy. I can imagine it’s not an easy business.

Looking forward, as CFO, I’m sure you’re thinking about how to invest the newly raised money and looking at the next three- to five-year horizon.

Where should we be looking? What is the company’s focus? What are the strategic pillars for you?

Bei: Great question. We think about this every day.

Going back to the basic AI paradigm, it’s always formed around compute, algorithms and data.

We’re really going to focus on those three things as we try to push the company to the next level.

Data is probably the hardest part because it’s not just about money. You need to think about smart ways of obtaining that data.

Talent retention and talent recruitment are also obviously the number-one priority for management.

You definitely know how expensive data scientists and algorithm scientists have become these days.

Grace Shao: How much are they making these days in China? Give us a range.

Bei: Not as much as their U.S. counterparts, I think. But even for college graduates fresh out of school, if you have the right experience and pedigree, you can make at least five to 10 times what a traditional software engineer might make.

It’s a very highly sought-after pool of talent.

Grace Shao: Okay, so are we talking about RMB 2 million to RMB 3 million? I’m trying to force you to give us a range. Five to 10 times is a big figure.

Bei: No, no. It’s a big figure, but software engineers don’t make as much as they used to anymore.

Grace Shao: The irony in all of this.

Bei: Exactly.

So talent is obviously a huge priority for us.

We’re also looking at interesting opportunities because we don’t know where the next technology is going to come from.

These days, acquisitions are really about people and talent.

If we see interesting algorithms or ideas coming out of labs, we’ll consider making our own moves.

As you know, we work very closely with Zhejiang University. We have a postdoctoral lab with Zhejiang University where we put a lot of effort into computer-vision research together.

Hopefully, we’ll identify talent and interesting early-stage products along the way.

Grace Shao: Would you go into hardware? Would you build your own robots?

Bei: Not right now. It’s already a pretty busy space.

But we’re definitely looking at hardware, probably not robots directly.

As I said, we’re already working with hardware companies to collect data.

We work with some scanner companies and LiDAR companies in China to collect 3D data.

To collect 3D data, you don’t only need cameras. You also need LiDAR, which gives you geometric information.

For example, we’re working with Hesai, which is a very good LiDAR company, to come up with solutions.

So we’re starting to dabble in hardware. We’re not purely a software company anymore.

Going forward, I think the two are going to be coupled together.

Especially in our world, if you want to have state-of-the-art 3D models, you will definitely need help from hardware companies, and we’ll probably do some of it ourselves.

Grace Shao: That makes a lot of sense.

You guys are well positioned given the amount of interest in you right now, and given that you’re in Zhejiang and close to Zhejiang University, where there’s a lot of talent coming out.

I think over the last year, the West has really opened its eyes to Zhejiang University, but in China everyone already knows it’s an absolute top-tier school.

You have people like Liang Wenfeng, and there’s just a lot of talent coming from that region.

Anyway, I really appreciate your time.

I want to ask you one last question, which is something I ask everyone who comes on the show: what is one differentiated view you hold? Something you think is non-consensus?

Bei: I think the TAM, the market for spatial intelligence, is actually going to be bigger than the market for large language models.

It’s still a little early, but if you look at human intelligence, language is only one part of our intelligence. Spatial understanding is also critical from an evolutionary perspective.

Eventually, if we can help AI crack that capability, the applications will be extremely broad.

I think there will be many things that robots or agents can eventually do that people probably haven’t even thought about yet.

It’s still early. It obviously requires a lot of exploration and effort, and there will be many pitfalls along the way.

But we believe this is a very, very large opportunity, and we’re fully committed to it.

That’s one view I think may still be a little bit off-consensus today.

Grace Shao: Thank you so much. I think we still have a long journey ahead.

Bei: Thank you, Grace.

Grace Shao: Thank you for your time, and congratulations again on the IPO.

Bei: Thank you very much, Grace. Nice talking to you.

AI Proem is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.



Get full access to AI Proem at aiproem.substack.com/subscribe

Podden och tillhörande omslagsbild på den här sidan tillhör Grace Shao. Innehållet i podden är skapat av Grace Shao och inte av, eller tillsammans med, Poddtoppen.