AI Proem Podcast
Avsnitt

From Beauty Apps to AI Agents: Meitu’s CFO Gary Ngan on the Future of Visual AI

Dela

In this episode, I spoke with Gary Ngan, CFO of Meitu, about how the company is evolving from its roots in consumer photo editing into a broader AI-native visual creation platform across photo, video, design, and agents. For many investors, Meitu is still associated with beauty editing and selfie apps, but Gary frames the company today as an AI application company serving both leisure use cases and productivity workflows.

We spent a lot of time on Meitu’s business edge: why visual AI is not just a foundation-model race, and why aesthetic judgment, controllability, and vertical context matter. Gary argues that visual creation is highly subjective. The same prompt can mean very different things across countries, cultures, product categories, and commercial goals. That is why Meitu is building verticalized products such as Picchi, DesignKit, Kaipai, Vmake, and RoboNeo, instead of relying only on one general-purpose AI model.

We also discussed the business model. Consumer subscriptions have become Meitu’s main revenue engine, while advertising is no longer the strategic growth driver it once was. Gary explained the shift in Chinese consumer willingness to pay for apps, the higher ARPU potential in overseas markets, and how new AI-native products like Picchi could introduce additional monetization through personalized models and AI credits. He also addressed AI compute cost, why more than 90% of Meitu’s AI outputs come from its own models, and why the company sees AI as a TAM-expanding opportunity rather than simply a margin risk.

Finally, we covered competition and globalization. Gary explained how Meitu thinks about competing with ByteDance, Kuaishou, Canva, Adobe, Shopify, Alibaba, and other AI-native visual tools, and why Meitu’s approach is more vertical-driven than general design-platform driven. Lastly, we touched localization, from different beauty preferences across markets to why true globalization requires understanding culture at a much deeper level than translation or marketing campaigns.

CHECK OUT THIS CONVERSATION. Gary’s so cool.

To find the previous episodes of Differentiated Understanding, see here.

Every episode, I bring in a guest with a unique point of view on a critical matter, phenomenon, or business trend—someone who can help us see things differently.

Season two will host a series of guests from analysts, VC investors, builders, researchers, founders, and product managers. For more information on the podcast series, see here.

Chapters:

00:00 What is Meitu today? Mapping Meitu’s product portfolio04:14 Why vertical focus still matters in the age of AI agents06:36 Aesthetic standards, subjective prompts, and visual AI nuance11:31 How AI changes art and creative expression15:02 MeituHub and MiracleVision as visual AI infrastructure17:01 Why Meitu needs its own models20:55 How Meitu chooses models and the role of designers25:09 Meitu’s AI legacy and generative AI strategy28:27 AI compute cost, ROI, and gross margin38:22 Subscription growth and advertising dependence40:47 Partnerships with consumer chatbots and platforms43:27 Competition with ByteDance, Kuaishou, Canva, Adobe, and others49:49 Deepfakes, misuse, and AI safety safeguards52:42 Globalization, localization, and cultural differences59:16 The biggest investor misconception about Meitu

Transcript (AI-generated, for reference only)

Grace Shao:Gary, thank you so much for joining us today.

Gary Ngan:Hi Grace. Good to be here.

Grace Shao:I’m really excited to have this conversation. To start, tell us what Meitu is up to these days. For a lot of investors and users, when they think of Meitu, they still think of the selfie and beauty-editing app. How would you define Meitu today? Is it still simply a consumer AI company, or is it much more than that now?

Gary Ngan:Meitu is no longer just a selfie or beauty-editing company. I would define Meitu today as an AI application company specializing in photo, video, and design.

We focus on very high-value verticals where we can leverage AI to deliver high-quality results to users. We often refer to these users as prosumers: people who have strong design needs, but no prior formal design training.

So that is how I would define Meitu today.

Grace Shao:That makes sense. Tell us about the products, because you have quite an array of them. Some are more consumer-facing, some are more prosumer-facing, and some may even be a bit more enterprise-facing. There is Meitu, BeautyCam, Wink, Picchi, DesignKit, Kaipai, Vmake, RoboNeo. Help us map out the ecosystem.

Gary Ngan:We think about Meitu’s product portfolio in two main categories: applications for leisure and applications for productivity.

Applications for leisure include the Meitu app, BeautyCam, Wink, and Picchi. They serve use cases such as photo-taking, photo editing, and video editing, usually for sharing on social media.

The Meitu app and BeautyCam are our core consumer applications. Wink extends our capability from photo to video editing. Picchi is our latest portrait-retouching agent, focused on personalizing editing styles.

The second bucket is applications for productivity, which includes DesignKit, Kaipai, Vmake, and RoboNeo. These products serve professional and commercial content creation needs.

DesignKit focuses on e-commerce product-listing design. It helps merchants and creators produce product images, model images, and marketing materials much more efficiently.

Kaipai and Vmake focus on talking-video and marketing-video creation. Kaipai is more focused on the domestic Chinese market in verticals such as insurance and real estate, while Vmake is seeing strong traction in the U.S. fitness and wellness market.

To give you a sense, as of May this year, Kaipai had about three million monthly active creators, and cumulative content creation exceeded 400 million pieces. For Vmake, ARR in the first quarter of 2026 was about US$4 million.

Then there is RoboNeo, our AI-native agent product launched in July 2025. It is currently targeting the AI short-drama vertical. Its agent workflows can support scriptwriting, characters, storyboards, visual generation, and asset management.

So that gives you a rough idea of the different vertical products. But the ecosystem logic is very important, because many new products come from user insights we observe in existing products.

For example, DesignKit came from the poster-design function within the Meitu app. Kaipai came from the AI teleprompter feature in BeautyCam. Picchi came from new user behaviors we observed in the Meitu app.

So our portfolio is not a random collection of apps. It is a structured expansion from consumer imaging into AI-native workflows across photo, video, and design.

Grace Shao:That makes a lot of sense. But in the age of AI agents, would it make sense for Meitu to consolidate a lot of these apps? Or do you still think it is better to keep them separate for different types of users and workflows?

Gary Ngan:In the age of agents, we still believe we should focus on high-value verticals, because different verticals have many differences.

First of all, aesthetic standards are very different. I’ll give you an example. The phrase “handsome guy” would be interpreted very differently in an application serving the U.S. market versus an Asian market.

Even within the Asian market, if you are addressing e-commerce merchants selling gym products versus formal apparel, the word “handsome” will also be interpreted very differently across those verticals.

So being able to separate these different verticals gives you a very good head start in focusing on the aesthetic standards that each vertical needs.

Also, users in different verticals have very different behaviors and workflows. It is very important to build those workflows and that know-how into each vertical in order to create the right products.

With agents, you can cover a slightly bigger boundary. But I still think you want to focus on different verticals to maximize the output for the user, and also make it more efficient and easier to market within each vertical.

Grace Shao:That is really interesting. You touched on something that a lot of people discuss when they think about visual AI, which is how to ensure consistency and accuracy when translating language into visuals, especially when text can be in different languages and words can be subjective.

As you said, if you say “handsome” and I say “handsome,” that could mean very different things in our heads. How do you ensure that identity, description, and nuance are not lost? You mentioned vertical focus, but what is the technical side of that?

Gary Ngan:Instead of calling it fragmentation, I would say vertical focus is very important. That sets the tone.

Behind that, we also have a large team of designers who control different points in the model fine-tuning process. They help set the right direction for the aesthetic standards within each vertical. That is something differentiated in our product offerings.

Then, if you move one step forward, the data flywheel is also very important. Users within a vertical give us data through their behavior: which photos they use, which photos they edit, which ones they throw away. That is very important for us to improve image creation.

We are also in the camp that believes controllability in visual applications should not just come from AI. You still need manual touch-ups at the end for users to make last-mile improvements, because aesthetic judgment is very subjective.

Even if an AI model works with you every day, you will always have subjective comments and small edits you want to make.

One other interesting point is that when we talk about aesthetic standards, in the case of leisure products, the face is usually yours. So you have a strong say and a strong sense of what is good for you.

The way we learn that is by studying the trend in your geographic location, giving recommendations, letting you try them, and then as you use the application, you tell us what is most suitable for you.

Picchi is a newly launched app where you can upload three to five sets of original photos and edited photos that you have done yourself. We are then able to learn that pattern and create a specialized model for you. The next time you want to edit a photo, you can call up the model that you trained yourself and apply your own aesthetic standard to your photos.

On the productivity side, however, aesthetic is not the ultimate holy grail of an image. It is very important, but whether that photo or video is effective in driving conversion, likes, or comments is also very important.

When we deliver images and videos to users, we take into account key data from that particular vertical and the metrics that matter for results. It is not just whether someone is subjectively handsome. It is whether this person, image, or video can help sell your product.

So the two camps are quite different.

Grace Shao:That is really interesting. From a pure consumer point of view, you pointed out an important nuance. If I upload my own face to a Meitu product, it might give me very smooth, pale skin and a more angular jawline or chin. But if I use an American fine-tuned product, it might give me more contouring. It is a very different aesthetic.

But to your point, if you are a prosumer, a content creator, influencer, or e-commerce seller, then it is not only about whether the image looks good. It is about whether it drives sales.

I want to ask something slightly more philosophical before getting into the businesses. Technology often changes art. Photography changed painting. Software like Final Cut Pro and Photoshop changed photography and video. How is AI now changing how visual artists approach their vision and craft?

I have also spoken to companies like Kuaishou, where they have Kling and are partnering with AI-native film studios. These people are creatives, but they do not view AI as disrupting their work. They use AI as a tool to create their work. What do you think about this at a high level?

Gary Ngan:If you look at our core value proposition, our mission statement is uniting art and technology.

One step down from that, we are trying to democratize design, art, and creative expression.

What AI changes is that it enables people who have creative ideas, but not the actual training or skills, to express those ideas.

A lot of the time, we have creative ideas that we want to express, but our motor skills are not refined, or we do not know how to put colors together. AI can help us deliver those ideas.

That is fundamentally changing the artistic landscape to a certain extent.

Other companies may say that existing professional filmmakers and designers can use AI to make things more efficient or create things in a different way. That is great. But I think the bigger impact on the world is enabling many people who previously could not create anything. They had ideas, but could not express them. Now they are able to express them.

That is what is fundamentally changing the industry.

Grace Shao:We have talked about how you have many different products, and you explained that they are targeted at different verticals. How should we understand MeituHub and MiracleVision? Are they the operating system underneath everything?

Gary Ngan:MiracleVision and MeituHub are the visual, image, and video infrastructure that we have. Our applications are built on top of these things, so they go hand in hand.

We are still fundamentally an AI application company, but we also need visual infrastructure.

To give you an example, over 90% of our AI outputs come from our own models. There are many situations where we think the models we fine-tune ourselves perform better than third-party models. There are things that other people do not necessarily focus on, so we have to invest in R&D and create that infrastructure ourselves.

MeituHub is also a way for us to export that technology. People can use our APIs and skills to build their own applications or integrate them into their own systems. That also reinforces our vision of democratizing design.

Grace Shao:That is a perfect segue to my next question. I understand your team fine-tunes your own models, but you also build on various open-source models. Why does Meitu need its own model?

Traditional application companies often did not need to own the foundation layer. So why does Meitu need that? And more broadly, why are so many Chinese consumer internet companies pushing out models? You see even companies in food delivery, ride-hailing, and other consumer internet sectors releasing models. Is this a cultural push, or something else?

First, how does Meitu think about it at the company level? And if you can comment, how do you view this competition across the China ecosystem?

Gary Ngan:It is harder to comment on the overall market, because what we do, visual image and video models, is quite different from language models. So I will focus on why we do our own models.

Our belief is that one general model will have difficulty performing well across all verticals, because context is so important.

If we do not have our own models, then aesthetic standards will be set by third parties. When a third party creates a model, they have their own idea of what aesthetic standards should be. They have their own idea of what should be generally good given a certain prompt word.

But that may not be applicable to the verticals we are working on. That is why we need our own models to serve those purposes.

At the same time, we integrate third-party models because even within a vertical, there are corner cases or edge cases that our core model may not be optimized for. In those situations, we call on third-party APIs to serve users.

As an AI application company, the only point of optimization is user satisfaction. We use a combination of our own models and third-party models to serve that purpose.

Sometimes we see more and more users calling third-party APIs for similar prompts or similar creation scenarios. Then we will augment our models to cover those scenarios as well.

Our model is continuously growing, but we make it very vertical-driven. We have told the market that we are not in the business of creating a general-purpose model. We are creating vertical models. But that does not mean we are giving up model training altogether.

Grace Shao:So there is a lot of industry know-how in each vertical that you have.

When it comes to which foundation model you choose for each task, how do you make that decision? I spoke to one of your colleagues at SuperAI, Rocky, your VP of R&D. We discussed the fact that you use a series of open-source models and also work with different model providers. What is the main factor in choosing which model to build on for which vertical? How do you delegate tasks across models?

Gary Ngan:At a high level, there are two main forces behind that.

One is user behavior. If a user uses Model A to create a certain task, and many users do not press save or do not continue working on it, then we probably need to serve that task with a different model. It is a data flywheel type of operation.

The other factor is our large team of designers, who are very involved in training these models. Designers help set the standard for what the right model should be for a particular task.

This is a very important differentiation for our company versus most technology companies.

I am not sure if you are aware, but our founder and CEO was an art student by training. In his day, he was the top student in the Tsinghua Arts Academy entrance exam for oil painting.

In his mind, aesthetic standards are always very important. Because of that, designers in our company have a very strong say in every product and every feature we launch.

Over more than a decade of designer training, the rest of the company has also developed stronger aesthetic standards. Product managers and R&D engineers also have quite high aesthetic standards now.

Our company is organized toward delivering the best aesthetic standards for users. That is very differentiated from most tech companies.

Most model companies may think: We solved this problem, the photo is done, the video is generated, the main character is stable throughout three minutes, so the mission is accomplished.

But for our designers, apart from the stability of the main character, they also look at whether the lighting is realistic, whether the color fits that vertical, and whether anything feels wrong from an artistic point of view.

Those are the things we really focus on when fine-tuning. That is something we are very proud of, and I think it is a major differentiator.

Grace Shao:Even as a consumer user, I can say your products have that extra last-mile touch-up tool that others often do not offer. It is meticulous and accurate. You can zoom into pores or details in the background. It is interesting to hear about your founder’s background and that artistic legacy, because that culture really shines through the products.

Speaking of legacy, I want to understand Meitu’s AI legacy and strategy. You have been working in image and video for over a decade, so you obviously have a vast database and deep know-how in visuals. How does that industry expertise translate in the age of generative AI and in the future agentic world?

Gary Ngan:Generative AI has changed the speed, scope, and value of what we can deliver.

First, speed. New AI capabilities can now be translated into user-facing features much faster, helping us launch popular effects globally and drive overseas growth.

Second, user experience. Generative AI enables effects that traditional computer vision technology could not fully achieve.

For example, facial and body retouching is no longer just manual adjustment. AI can reconstruct details, lighting, and texture in a much more natural way.

Third, target addressable market expansion. AI helps us broaden into productivity workflows like DesignKit and Kaipai, which were things we traditionally could not do.

Overall, AI is very empowering in helping us get to where we want to go.

Before AI, all we could deliver was better tools. But in order to use those tools, you still needed pretty good aesthetic standards or some understanding of the basics. Otherwise, giving you those tools did not really help much.

With AI, you still need maybe 10% or 20% of that understanding, but the requirement is reduced massively. AI can give you many choices to choose from, and then you can start building from there.

That helps us move from leisure applications to productivity applications. That is really what the strategy is about today.

Grace Shao:AI can act like a guide or mentor if you are new to a certain craft or sector.

Let me ask the spicy question. AI compute cost is obviously extremely high. Image and visual generation are expensive. How does the economics work right now? Does AI compute affect your gross margins, or are you seeing ROI already?

Gary Ngan:As I said, currently over 90% of our generative outputs come from our own models. As long as we are using our own models, the cost is very manageable. Our gross margin is still over 70%.

Also, when you are editing your own face or editing a product photo, these things are not purely AI-generated. You may want AI to edit a little bit, remove someone from the background, or create a new background for a product, but the entire photo is not purely AI-generated.

It is true that AI inference has a cost, but it is not as if every photo now incurs a lot of cost. We need to make that distinction first.

As we move into new verticals, like music videos and AI short dramas with RoboNeo, those are more experimental. We are using more third-party models, so margins on those new applications will be much lower than something like Meitu Xiuxiu.

But as we continue to progress, we will develop our own models to replace some of the third-party costs. Over the longer term, we also believe API costs will come down.

So we do not see this as a threat. In fact, the integration of AI has expanded the addressable market so much that it is a much bigger opportunity than threat.

Grace Shao:I appreciate that nuance. You are explaining that the first type of usage does not use as much AI or token cost as people might expect from the headlines. The second part may be more expensive, but we are still in very early stages.

Let’s take a step back. For some of our American or Western audience, they may not be as familiar with Meitu. How do you fundamentally make money?

In your public disclosures, consumer paid subscribers grew more than 30% year-on-year. What is driving that growth? Is it that the AI features are much better now? Is it global expansion? Help us understand the business model and what is driving growth.

Gary Ngan:Our main revenue source is subscriptions, mostly on the leisure side.

That is our second growth curve. The first one was advertising, but that business has matured.

The second growth curve, which is still growing quickly, is subscription on the leisure side. The main driver has several parts.

The first is China user behavior. Paying for apps really started after COVID. Before COVID, virtually all applications were free. They competed through free usage, advertising, or redirecting traffic to other applications to generate money.

After COVID, many user-facing applications realized advertising was under pressure, and they wanted new revenue sources. Without colluding, many of them started charging users. That kickstarted the user subscription process.

What is less understood is that users then began realizing that applications have to be paid for. As time goes by, the behavior of paying for applications grew on them.

Now there is much less of the issue of, “This app has to be paid, so I am not using it.” That was a real mindset before. Now it is more like, “This app costs 15 RMB a month. Is it worth it?”

That is what I would describe as the beta factor, meaning the overall market. Users are becoming more and more used to paying for mobile products.

That is one reason we are confident that paying subscribers and the paying subscription rate of our leisure applications can continue to grow.

To give you a sense, we have done surveys. The global paying percentage for photo and video applications is about 20%. If you benchmark music and video apps globally versus Chinese equivalents, the Chinese equivalent is usually around half. For example, if Spotify is around 40%, the Chinese equivalent might be around 20%.

So if global photo and video applications are at about 20%, China should at least achieve about 10%. Right now, we are around 5% to 6%. So there is still another 80% to 100% growth headroom there.

The second growth potential is international expansion. In high-ARPU areas like the U.S., Europe, and East Asia, including Japan and Korea, the base ARPU is already much higher than China, anywhere from 100% to 200% higher. The paying percentage can also be much higher.

To give you a sense, one of our applications called AirBrush has over 50% paying percentage in the U.S.

As we launch stronger operations in these high-ARPU countries, we expect our blended paying percentage to grow further.

One final point about monetization is that we are integrating more generative AI capabilities into these applications. For example, Picchi is an application for leisure, but it uses an agent for editing, and that has a completely new business model.

On top of regular subscription, if you want to create your own model to apply your own editing skills, you have to pay for that model separately. That is another monetization test we are currently working on.

Grace Shao:When I was reading your earnings reports, I was a bit surprised that your highest revenue generator is consumer subscription, because the default mindset is that people have very little willingness to pay.

But as you said, whether it is the change in behavior in China, or people having more appetite for premium add-ons or AI-plus features, willingness to pay is changing.

There is also the fact that advertising can be annoying to sit through. You do have a lot of advertising, I have to say. Spotify does too, and I think that drives people to pay to get rid of advertising.

On that note, do you think you will gradually reduce your dependency on advertising? It is still your second-largest revenue model.

Gary Ngan:We have not relied on advertising since 2022. At the corporate level, we made the point that we are no longer strategically trying to drive advertising.

You have seen our advertising business grow at low single digits over the past few years. Advertising is not what we are fundamentally trying to drive.

However, we are experimenting with advertisers on fun and engaging AI-infused campaigns.

It is hard to describe with words, but you can imagine users generating viral photos with a brand advertiser’s branding that fits the brand image. That gives the uploader a lot of likes and gives the advertiser a lot of exposure.

So we continue to experiment with those things. But in any case, we are not relying on advertising for business growth.

Grace Shao:On partnerships, I had this idea and I do not know if you are doing anything like this. Would you partner with some consumer-facing chatbots in China to help them with video and visual capabilities?

For example, could someone go into a consumer chatbot and call up Meitu’s capabilities? There may also be competition there. How do you view your relationship with these players?

Gary Ngan:We are open. In fact, we are already an official partner with WeChat, not on the Xiaochang side, but in another area. I do not remember the exact English name, but basically when you are using the chatbot, you can call up Meitu.

Right now, it is still a lighter relationship, almost like traffic redirection. But our goal is to democratize design. Being able to work with more people and enable more people to access that power to express themselves is something we are open to.

Grace Shao:That makes sense. It feels like they may not want to put as many resources into this specific use case, and you have the know-how in doing the best video and image editing.

Gary Ngan:I would not say they do not have the edge. I think they may just not want to focus on that.

Creating these applications requires a lot of focus. It requires the right organizational structure and a laser-sharp focus on trial and error, and on creating the best aesthetic output for users.

These may not be the things that larger companies want to invest in. It is important relative to our size, but to them it may be something they do not want to focus on. If they wanted to do it, I think they could.

Grace Shao:Let me challenge you a little bit on big tech. In China, ByteDance and Kuaishou clearly have a lot of edge and moat from massive pools of image, visual, and video data. In the West, we have Canva, Adobe, and other global applications. Even Shopify and Alibaba are creating e-commerce staging and design tools.

In this big world of competition, or peers if we put it more nicely, how do you see Meitu’s strength? Who are the most relevant competitors that are similar to what you do? And who may look similar on the surface but are not actually doing the same thing?

Gary Ngan:We have to separate it into two categories.

On the leisure editing side, with the exception of one business unit within ByteDance, there are not many large companies doing that globally. I do not think there is any real large company doing that in the U.S.

There are smaller companies, but they are much smaller compared to us.

On that side, our edge is really continuing to follow and set the trend for the latest aesthetic standards and what helps users stand out on social media. These are the things we have been doing for more than a decade, and we will continue to excel in them.

On the productivity side, there are many companies doing similar things, but taking a much more general approach.

For example, Canva and Adobe use one product to satisfy different verticals. Adobe is organized around media: photos, vector diagrams, video, effects, and so on. Canva is one editor trying to fit many situations. Figma is also one app serving many applications.

They are design-driven. We, on the other hand, are much more vertical-driven.

We are not restricting ourselves to a specific media type. We are saying, within e-commerce, what do you need?

You need product photos. You need very short product videos. You need the ability to generate batches and batches of photos. You need to know the latest trend on the e-commerce platform you are selling on. For that particular product, you need to know the selling points. You also need to know the rules of Amazon or Temu and what you need to abide by when selling those products with pictures.

All these things are baked into DesignKit.

If you are using Canva, I highly doubt it will have a red flag saying, “You should not be using minors in this product photo.” Canva may not even know that you are creating a product photo in the first place.

So these are the different focuses we have.

In terms of competitors, it is hard to say who is a direct competitor, because at the end of the day, you can use Photoshop, Canva, or our products to create an e-commerce photo. They are all peers, but we take different approaches.

If we take a step back, generative AI is still very early. It is 2026 now, but generative AI really only started in earnest late last year for visual use cases. Before then, a lot of generative AI photos still looked AI-generated.

Grace Shao:They were quite bad. There might be six toes, or the face was disproportionate.

Gary Ngan:Even if there was nothing obviously wrong, you would look at the photo and know it was AI-generated. It did not feel real.

Now we are just starting to see things that are harder to distinguish between human-made and AI-made. This is how we can empower the industry and increase efficiency.

We are still very early in this market. That is why we are very optimistic and see a lot of opportunities.

Grace Shao:A little side note: in 2019, when I was still with CNBC, I covered deepfakes. At the time, there were a lot of deepfake videos of Obama or Zuckerberg. A startup even made a deepfake of me. It was literally just plugging someone else’s face onto my head and body, and nothing really worked.

But now, fake images and videos are getting very hard to distinguish with human eyes. How do you view the ethical side? How do you stop misuse of the technology?

Gary Ngan:First, we have put in safeguards.

For example, on Picchi, if you generate a model of yourself using your own photos, that model cannot be applied to anything other than your face. If we detect that it is not your face, we will not allow you to apply that model to another face.

In some of our generation applications, we have also put in safeguards around certain words, such as violence or pornographic images. You cannot generate those using our applications.

So there are safeguards that we put in place. Obviously, we can only do so much.

One thing that makes it slightly easier for us is that we organize our applications into different verticals. Users come into our applications with a very strong intent. They know they are creating e-commerce photos, for example.

Instead of giving them a general chatbot where any random person can come up with a random idea like putting their face onto the President of the United States, it is harder to imagine someone using DesignKit to run a prompt like that.

Organizing into different verticals also helps us mitigate the risk a little bit.

Grace Shao:The last area I want to talk about is globalization and your global strategy. Meitu is globally available. It is interesting because, as you said, you focus on each vertical, and you have not done a big splashy general marketing push. It also feels like that is true geographically. You are in Southeast Asia, Japan, Korea, Europe, the U.S., and so on.

Help us understand global scaling. What have been the challenges? How have you done it successfully? And how do international users from different regions behave differently from users in China?

Gary Ngan:I will answer the second part first. Users in different regions all behave very differently.

Grace Shao:Give me all the stereotypes.

Gary Ngan:Not stereotypes, but I will give you one example.

We were doing a user focus group in the UK and spoke to a male influencer. He said, “Your app can edit my jawline? That is incredible. I would totally pay for it. But I do not think it is a good idea to smooth out my skin.”

Grace Shao:That is interesting. So it is not okay to pretend you have better skin, but it is totally okay to have a chiseled jawline?

Gary Ngan:He did not mention whether it was ethical or not. That was just his feedback, word for word.

The challenge, or the interesting thing, is that we have to really listen to what users want in those markets. Different geographic locations need the right mix of features and marketing campaigns.

I will give you another simple example. A few years ago, we were looking at Lunar New Year. Koreans also celebrate Lunar New Year, and in China Lunar New Year is a festival where we get a lot of usage.

We had launched features in China that were very popular that year, but in Korea there was no uptick. Later, when we had local Korean colleagues helping us run marketing campaigns there, they told us that Koreans generally celebrate Lunar New Year with white clothing and a white theme, while Chinese people celebrate with red.

Our Chinese marketing team was surprised, because in China, white is usually associated with funerals. It did not register.

That example tells us there are many things we need to immerse ourselves in culturally to understand how people behave, what they care about, and what the standards are.

We cannot stereotype anything. Every place and every person behaves very differently.

That is the biggest challenge, but also the biggest opportunity.

Now we are setting up offices in different parts of the world. We are sending product managers overseas regularly to do more focus groups and, more importantly, to experience the lives of the users they are trying to serve.

In the past, we relied too much on consultants, reports, or reading online. That is not enough anymore. We are fixing that, and I think we are making progress in some countries.

Fingers crossed, we will continue to grow bigger in Western markets.

One other tailwind that has helped us is TikTok and K-pop. Back in the day, editing a photo seemed socially unacceptable to a certain extent. But with TikTok, people are more relaxed about filters being applied and playing around with your face. It is no longer as taboo in many Western countries.

The rise of K-pop is also influencing cosmetic styles, and that becomes a segue for us to try different things in Western markets.

There are very interesting things happening. But the most important thing is for us to really understand, live, and breathe those cultures so we can create things users want.

Grace Shao:That is meaningful, and it feels important for a new generation of Chinese companies going global. Localization cannot just be reading headlines or high-level reports. You have to understand the culture, because culture influences the business.

To wrap up, I really appreciate your time. My last two questions: first, what is the biggest misconception investors currently have about Meitu’s business?

Gary Ngan:One of the biggest misconceptions is that general models are going to destroy everything, and that there is no place for AI applications.

We think that is quite unlikely on the visual side. I am not sure about the language side, but on the visual side, aesthetic standards are very subjective and personalized, and a lot of controllability is needed.

Different verticals have different interpretations of the same words. So the way models are trained and organized, even with agents, makes it unlikely that a one-size-fits-all general model can satisfy all verticals.

Every vertical has its own workflow and standards. AI application companies are very important in making those adjustments and optimizing workflows for users.

That is the biggest misconception.

Grace Shao:I agree with that. We are seeing more of that realization in the market now. You have strong vertical use-case AI-native companies coming through, like Harvey. I have also met companies where former investors are building equity analyst research tools.

You can say generic GPT can be used for research very easily. But to your point, these teams know the niche use case. They know the process, the standard, and the workflow better than anyone else. Even if the TAM is small, it can be big enough for their business.

The last question I ask everyone on the podcast is: what is one differentiated view you hold? Something that is a bit against consensus.

Gary Ngan:Is it related to the company or the industry?

Grace Shao:It could be anything. Usually people answer about their topic, but it can be anything.

Gary Ngan:I think life expectancy will be a lot longer than we think today for our generation.

Grace Shao:So we are going to live to 150, thanks to Bryan Johnson’s experiments?

Gary Ngan:Possibly. Then there will be more time.

There is a lot of advancement in AI. It speeds up many pharmaceutical processes. You can run different trials much faster and understand the underlying issues more efficiently than before.

And with more time, there is more time for us to create more art.

Grace Shao:And live a healthier life. Although right now, anyone working in AI knows AI never sleeps, and I think we are all working more than ever.

But thank you so much. That is definitely a differentiated view. I really appreciate your insights and your sharing today. Thanks again, Gary.

Gary Ngan:Thank you so much, Grace, for this opportunity. Really nice talking to you.

AI Proem is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.



Get full access to AI Proem at aiproem.substack.com/subscribe

Podden och tillhörande omslagsbild på den här sidan tillhör Grace Shao. Innehållet i podden är skapat av Grace Shao och inte av, eller tillsammans med, Poddtoppen.