Story

A Decade of AI Talent Migration in China

A ten-year migration links China’s first computer-vision startups to today’s foundation-model teams, carrying engineering knowledge and the memory of failure into a new generation.

Chinese AI researchers and engineers connected across laboratories, startups, and cities

In 2026, Moonshot AI released the open-source Kimi K3. On several benchmarks, it approached or surpassed the leading closed models of the time. It also created a new “Kimi moment” in Silicon Valley.

The impact was different from the “DeepSeek moment” a year earlier. Until then, the most common descriptions of Chinese models abroad were cheap, open, and good value. K3 suggested another possibility: a Chinese team could push toward the frontier of model capability and still choose to open-source the result.

The model arrived in 2026. The people who built it did not appear overnight.

Over the past decade, I have worked at a computer-vision company, an AI institute, and an investment firm. I was also involved in Moonshot AI’s early financing. Some of the researchers and engineers I knew moved from Microsoft Research Asia, Baidu Research, and Chinese university labs into the first generation of AI startups. A few years later, they moved again—from SenseTime, Megvii, Yitu, and their peers into large internet companies, quantitative trading firms, and a new generation of foundation-model teams.

From the outside, this can look like two separate groups of founders taking the stage one after another. From within the industry, it looks more like a migration that has lasted ten years. What moved was not only talent, but also faith in technology, habits of engineering, and the caution left behind when an earlier idealism collapsed.

What the first generation taught its people

I joined Yitu in 2016. Together with SenseTime, Megvii, and CloudWalk, it belonged to a group the Chinese media called the country’s “four computer-vision unicorns.” All four were trying to move deep learning out of papers and test sets and into cities, hospitals, and other real-world settings.

On my second day, a founder called the three employees who had joined that week into a small meeting room and announced that, as of that day, the product team existed. Each of us received a product line and was told which researchers and engineers we would work with. Then we were left to it.

Product managers had no clearly defined remit. We did anything the researchers and engineers did not do. Nobody could tell us what the standard process was because the company itself did not know.

The first year was painful. Every process began with a mistake. A project failed, so we learned where to add a check. A collaboration went off course, so we added another review the next time. New employees at a mature internet company could learn an established method. We had to collide with a problem before we could work out how to solve it.

Two years later, I had become grateful for that experience. When a process tells you what to do, it is easy to stop asking why. Every step we created came from concrete feedback, so we knew which problem it was meant to solve. That way of learning later transferred to completely different industries and roles.

The first-generation AI companies were also willing to give very large problems to very young people. Their teams included undergraduates, students who competed in programming contests, and even high-school interns who had entered research labs early. The founders did not have the right answers. They simply believed that smart, curious people should be allowed to try.

Today’s foundation-model companies are often described as a new kind of organization. To those who lived through the previous cycle, their character is familiar: managers set a broad technical direction, while researchers and engineers explore with considerable freedom. Many systems are not designed in advance. They grow gradually through failed experiments, system crashes, and the work of turning models into products.

The first generation also faced very concrete limits. A facial-recognition model might perform well in one province and fail in another because of differences in climate, light, camera height, or angle. The team would have to collect and label new data, then train the model again.

The more customers a company acquired, the more researchers and delivery staff it needed. Unlike internet software, these systems did not keep reducing marginal costs as they scaled. Some of the most expensive science and engineering graduates in the country ended up building custom solutions for one client after another. Revenue grew, but costs and organizational complexity grew with it.

This was the shared constraint of the first generation: when models could not generalize, the business began to resemble high-end outsourcing.

Those companies did not become what people had originally imagined. But they trained a generation. Their employees learned how to work with data, turn papers into systems, and understand the distance between a technical metric and a real deployment. They also lived through a funding boom, rapid project expansion, and an industry downturn. When the second AI wave began, they carried both the experience and the disappointment with them.

The talent networks came first

This chain of talent reaches further back.

Before the four computer-vision unicorns existed, Microsoft Research Asia (MSRA), Baidu Research, and university labs had already trained a generation of Chinese AI researchers. Among my university classmates, many of the strongest computer-science students interned at MSRA. A similar path was common among people who later founded or joined the first generation of AI companies.

Those companies then trained the next generation. Many founders and key employees at Moonshot AI, MiniMax, Zhipu AI, and other foundation-model teams once worked or interned at first-generation AI companies. They may not have worked together at the same time, but they share a piece of industry memory.

Universities and programming competitions formed another layer of the network. Labs, elite computer-science programs, and competition teams at Tsinghua University, Peking University, Shanghai Jiao Tong University, Zhejiang University, and other schools let people learn very early how their peers approached a problem. When startups were formed, founders usually began with people they had already worked and argued with.

This was not a talent map designed in advance. When people decide whether to investigate a problem with no clear payoff, trust often comes from prior contact. You know what the other person is genuinely interested in, and you have some sense of how they respond when an experiment fails or resources run short.

That is why the composition of different companies diverged. Moonshot AI and MiniMax had more early connections to the previous generation of AI companies. Zhipu AI grew out of a Tsinghua lab and had easier access to students in the same academic network. High-Flyer, the quantitative trading firm behind DeepSeek, was more familiar with recent graduates from algorithm competitions and young engineers.

When I looked for Chinese teams capable of building general-purpose foundation models in the first half of 2023, I did not notice DeepSeek. My information came mainly from the previous generation of AI practitioners, while DeepSeek was recruiting younger people. I began to see it only after getting to know a new cohort of researchers.

This does not necessarily mean that the companies had fundamentally different definitions of excellent talent. The origins of the founding team determine which people it can reach first and most efficiently. The team that eventually forms is partly the result of the founders’ existing social radius.

Would people return after failing once?

ChatGPT was released at the end of 2022. After using it for a week, my colleagues and I knew that something had changed.

Earlier natural-language systems usually solved specific tasks. ChatGPT showed much stronger generalization. It also made the decoder-only path toward artificial general intelligence, or AGI, feel tangible for the first time. Over the next two or three months, we spoke with almost every researcher we could find.

We were not asking who planned to start a company. We wanted to know who had believed in scaling models before ChatGPT appeared, and who was willing to devote years to a question with no visible return.

Yang Zhilin quickly appeared on the list. He had worked on Transformer-XL and XLNet and had long been part of China’s early foundation-model research community. My own network was in computer vision, with fewer connections to natural-language researchers. It took us about four months to meet him. I visited Moonshot AI’s other founding members and repeatedly tried to add him on WeChat, without success.

At the same time, Moonshot AI’s early team needed more systems and infrastructure engineers. I began calling friends from the previous generation of AI companies. Some were still at their old firms, some had joined large internet companies, and others were working in quantitative finance. I asked whether they wanted to return and build foundation models.

Roughly 80 percent said no.

They had already given AI the most energetic years between their twenties and thirties. The first wave had not produced the companies they imagined, nor had it delivered the financial outcome everyone hoped for. After living through a complete cycle, waiting on the sidelines was a reasonable choice.

A small minority still wanted to try again. They did not begin by calculating titles and compensation. They simply thought the problem was worth joining. After several engineers joined Moonshot AI, they mentioned to Yang the investor who had been trying to meet him. We were finally added to the same WeChat group and soon flew to Beijing.

What struck me most when I first met Yang was his lack of affectation. He discussed technical questions directly and did not make the work sound mysterious. Later, after the first Kimi model and product were released, I asked the team how they had achieved the result. There was no secret, they said. They had simply done every necessary part carefully.

At that 2023 meeting, Yang used “LTV” to summarize three areas he cared about: long context; truthfulness, or reliability; and video, including multimodal data more broadly. Long context would let a model work with more information. Truthfulness addressed hallucination. Multimodal data could supply capabilities beyond text.

In business, LTV usually means lifetime value. In that conversation, the same initials became three technical problems.

Moonshot AI’s Series A financing did not rest on a conviction that the company was certain to succeed. The larger question was whether a startup had any right to enter the race. Under scaling laws, model capability grows with data and compute, while the required investment may have no clear ceiling. Large companies have cash flow, users, and infrastructure. A startup struggles to fund both a frontier model and a mass-market application.

There was no settled answer inside Meituan. Near the end of the discussion, Wang Xing said: “I do think the probability that a startup can build both a super model and a super app is low, but I am willing to back the attempt.”

That sentence helped move the investment forward. The wager was not on a proven business model. It was on whether a group of people deserved the chance to keep researching.

A failed Chinese GPT-2 experiment

Having talent and resources does not mean a team will understand a technological shift. In 2019, I was working at the Sinovation Ventures AI Institute. At one point it employed more than 200 people. Its job was to follow new academic breakthroughs, build product prototypes, and decide whether any could become a new company.

That year, we decided to train a Chinese-language GPT-2.

The conditions looked favorable. While reorganizing a data center, an investor in the Middle East found several hundred unused NVIDIA V100 GPUs and asked whether Sinovation Ventures wanted them. The team already included a group of natural-language researchers. We appeared to have compute, talent, and Chinese data.

Two months later, the model was ready. I asked it to write a simple email. It produced three commas in a row.

We were disappointed and initially blamed the quality of Chinese-language data. Looking back several years later, I think the main problem was ours. We did not truly understand scaling laws, and we underestimated data engineering. The paper described the model architecture, but training quality depended on collecting, cleaning, and organizing large amounts of data. The team consisted mainly of researchers, and we did not put enough people into doing that work well.

We tried again in 2020. The new team included more people with deployment and engineering experience and paid greater attention to data. It eventually incubated Langboat. But our way of thinking still belonged to the first AI generation: find an enterprise vertical, then sell model capability into it.

By then, I had developed a dependence on the lessons of failure. The first startup wave had taught me that selling model capability was not a good business. That conclusion was not entirely wrong, but it kept me from seeing soon enough that a general-purpose model would change the question itself.

Only after ChatGPT appeared did I realize that in 2019 we had come close to a direction that would later transform the industry—and that we had possessed an unusually large amount of compute for the time. What stopped us was not just a lack of technical ability. We were still using the product and business frameworks of the previous generation to explain something new.

Experience can be inherited. It can also become a burden. That is an easily overlooked part of talent migration.

Scarcity changed how Chinese teams worked

From 2023 through 2024, Chinese foundation-model startups generally lacked compute. Some North American teams already had single clusters with tens of thousands of GPUs. Few Chinese companies could reach the same scale. Building a general model requires continuous training, while serving users creates an ongoing inference bill.

Competition for consumers was especially intense. Products such as ByteDance’s Doubao and Alibaba’s Qwen had large internet companies behind them, and Chinese users had not formed a habit of paying a US$20 monthly subscription. Startups competed with free products while trying to convince investors that they still deserved more resources.

Kimi experienced this tension in 2024. Its long-context capability produced a period of organic growth, after which Moonshot AI increased advertising and tried to acquire more users. From the outside, those moves could look like a retreat from AGI. I see them more as survival under resource constraints. More users brought new tasks and feedback, and growth helped persuade those who controlled resources to keep investing.

The danger is that short-term metrics can reshape the research goal. Model capability, user growth, revenue, and financing do not always point in the same direction. A startup must keep deciding which actions are building the conditions for its next stage and which are pulling it further away from the original question.

Limited compute also forced Chinese teams to optimize training and inference. DeepSeek, Moonshot AI, and Zhipu AI all developed strong engineering capabilities under that pressure. Low cost was not an incidental by-product. It let more people use the models and gave open-source communities room to keep building.

Yet “value for money” also constrained what outsiders imagined Chinese models could be. A cheap model was easy to treat as a discount substitute for an American product. That began to change when DeepSeek-R1 appeared in early 2025. It let technical and general users see a model’s reasoning process and showed that improvements to both model and product could translate directly into adoption.

R1 also changed the capital market. Model companies could spend less time answering “When will you make money?” and direct resources back toward research. Kimi K2, Kimi K2.5, and Kimi K3 followed. As reinforcement learning and test-time compute expanded, the researcher’s ability to define tasks, design rewards, and judge answer quality once again shaped the model’s ceiling.

What excites me most about K3 is that a Chinese team is no longer competing on price alone. It approached the strongest closed models of its time, then released the result as open source.

A company’s environment determines whom it can recruit

Talent migration is not the movement of résumés from one company to another. The same person may perform very differently in a different environment.

The first-generation AI companies had a rare openness. The problems were new and managers did not know the answers, so young people had substantial room to explore. General-purpose models still require that kind of environment. When management drifts too far from research, the reward signals received by the team can gradually shift toward revenue, scale, and internal process.

Moonshot AI’s office has a white piano and Pink Floyd records. Its meeting rooms are named after rock bands. Those objects do not make a model more capable, but they reflect the same preference that allows a team to investigate things that are “not useful yet.” The outcome of scaling cannot be fully predicted. A company can only let excellent people keep experimenting and make sure they receive reliable feedback quickly.

Large companies provide a different environment.

ByteDance can connect model-generated content, consumption, and platform distribution. A video model produces content, users watch and respond, and platforms such as Douyin distribute it. That is a systems capability built during the mobile-internet era.

Alibaba trains its own models, invests in several startups, and supplies cloud infrastructure and compute. It plays part of an infrastructure role, allowing more teams to enter the competition.

Tencent’s models may not rank in the first tier at every stage, but WeChat is already embedded in the communication, work, and service routines of hundreds of millions of people. That environment is difficult to reproduce. If models enter it, interactions between people and AI will generate enormous amounts of context and feedback.

A startup can redesign its organization around a new problem. A large company already has a real, high-frequency environment in which people act. Neither simply replaces the other.

The strongest model only buys a company time

I still hold a controversial view: in the long run, selling model capability on its own will struggle to sustain high margins.

During periods of rapid progress, the leading model does have pricing power. When only one company can provide a capability, it can earn extraordinary returns. The problem is that a technical lead rarely lasts. Competitors narrow the gap, API prices fall, and foundation models gradually come to resemble standardized infrastructure.

A leading model buys a company time. The company must use that time to build an environment in which interactions between users and AI continually create data and feedback, allowing model and product to improve in return. Once that environment exists, more tasks take place inside it and later competitors cannot easily copy the whole loop. Without it, the original advantage disappears when others catch up with the model.

AI coding tools already offer an example. A good coding product is not a model placed behind an input box. It knows the context of the codebase, can call tools, and can run and debug programs. A programmer immediately accepts, modifies, or rejects a proposed solution. Those actions create clear reward signals.

The “super app” I expect looks something like this: people and AI continually produce new capabilities through real work. A model API alone is not enough.

People are also part of the environment when companies adopt AI. Knowledge workers may worry that once they teach AI how they work, their jobs will disappear. A forward-deployed engineer (FDE) must understand not only the technology, but also the organization and the anxieties attached to particular roles. That is labor-intensive work, but it may be a necessary stage in bringing AI into real companies.

The migration continues

At the end of 2025, I left investment and began building my own AI product. The decision grew out of a frustration that had lasted for years: since I first used ChatGPT, I have often felt like a “second-class citizen” in the AI world.

People who can write prompts, build agent loops, and draw on engineering experience can make models perform complex work reliably. An ordinary user meets the same empty input box and does not know how to describe the task, inspect the result, or repair a mistake. Even when AI builds software for me, the code is often impossible to maintain.

If a consumer product makes part of its audience feel unintelligent for years, the user cannot carry all the blame.

I am now trying to build a tool for communication between people. Human communication is not one person encoding a message and another decoding it. The participants create a shared field, gradually approaching one another’s intentions through response and correction. AI can take part in that process, but it should not require everyone to become a prompt engineer first.

That is why, after watching a decade of migration, I became a founder again. The first AI companies taught me how to work without an answer, and made me wary of businesses built on selling model capability. Foundation models forced me to examine those lessons again: which still hold, and which are simply habits left by the last failure?

China’s AI companies have now changed generations twice. Many researchers and engineers never left the underlying problem. They changed institutions, roles, and tools. People trained by early research institutes moved into the first startup wave. Some later joined large internet companies or quantitative firms; others returned years afterward to foundation-model teams. A younger generation arrived from universities and algorithm competitions.

This migration does not guarantee success. It simply means the next wave does not start from zero. Engineering methods from the last cycle can be reused. The memory of failure keeps reminding people that technical leads disappear, resources influence direction, and organizations can push smart people toward very different outcomes.

I do not know which company, if any, will build AGI, or what kind of product will bring it into ordinary life. But over the past decade, even when companies closed, changed direction, or lost their place at the center of the industry, researchers and engineers kept moving. They carried what they had learned to the next stop. They also carried the things they had misunderstood—and judged them again.

Sources and checks

  1. Original interview transcript provided by the author
Portrait of Ge Xu

About the author

Ge Xu

A coherentist

Ge Xu
A Decade of AI Talent Migration in China | China, in Fact