I quit my corporate job at Microsoft and my academic job at MIT this Friday. I still remember all the details of applying to these jobs in Spring 2026, vividly. It was a painful and uneasy process, especially when you know you will get a lot of “No”s. And right now, I have decided to start something new. I call it the biggest bet (I know we are building a simulation-based evaluation company… in the future, maybe there will be more biggest bets, but for now, it’s the biggest).
I want to leverage this post to explain: Why now? How? What?
Did you tell your mom?
When I told my colleagues and managers about this decision earlier this week, someone asked, “Did you tell your mom?”
I laughed. First, I think I am a mature adult and can be responsible for my decisions; second, I have never felt this passionate and excited.
Yes, indeed, they have known since this May, during my graduation trip. In my family, if it’s a No, they will tell me right away; if it’s a Yes, they will never respond. Perhaps this is my family’s philosophy?
Why now?
As I always think things have two dimensions, I will also ask: Why not now? There are a lot of opportunities in this AI era, but opportunities pass so fast. I have rarely felt “opportunity” so real. Back in the summer of 2022, when I worked on my first HCI paper, I first heard about a language parrot called Chat GPT 3. It was in a Google Colab and everything was code. The generated content seemed interesting and, potentially, revolutionary.
Then it became OpenAI. Then this “simple” next-token prediction spread from chatbots to coding, reasoning, audio, image, and video.
MatrAIx started with my long-term research collaborator Dr. Xiaomin Li and me. We started by curating a benchmark for AI in healthcare, then activation steering to mitigate safety issues, then skill induction to make workflows efficient and easy, and now something new. We are still young (fortunately under 30, and mentally and physically we are okay with failures and rejections) and can still afford the price of losing.
We even joked that we are “looking forward to failing” so we can quickly learn the winning paths (and all roads lead to Rome!). And that’s exactly what MatrAIx is looking for: looking to fail during pre-screening, then figuring out all the diverse ways to fail, so that the only routes left are to win and succeed.
How can MatrAIx solve this problem?
The second half of the AI era is coming. It feels real. The first half, the “training season,” is almost over and saturated, and chasing the hardest problems already proves to be feasible; then the second half of AI is coming: Evaluation and Verification.
How can we apply AI or LLMs to real-world people, solving their real needs and real tasks? How can we find the edge cases and failures in real-world deployment? How can we close the gap between good AI and bad user experience? The AI can be extraordinarily smart and capable, but human users might not have smart questions all the time. So how can we bridge the gap in this type of “not so hard” long-horizon task? How can we keep AI capable in multi-turn conversations instead of catastrophic forgetting?
Then, it’s time for MatrAIx.
What are you going to do with MatrAIx?
MatrAIx has a strong foundation in the open-source research community. From 2 people discussing on Zoom in mid-June, to 10 friends who all shared the same passion that “the AI evaluation era is COMING,” to 100 friends from academia and industry who echoed it and contributed to building personas, environments, and applications, to now, when our Discord has 500+ community contributors working on layers in applied AI, evaluation benchmarks, and simulation-based model training mechanisms.
This week, our core founding team members are flying from the East Coast to the West Coast and firmly believe in this ambitious goal: creating the evaluation infrastructure for simulation-based agents. Vision, mission, and imagination are the core of a successful team.
To be honest, some VCs don’t share the same vision. We felt sorry about that. The understanding of the FUTURE is usually in a few people’s hands. I have always found that the people who say “but…. it’s too early” are usually the people who say “oh no! it’s too late”. So the conclusion is: don’t convince and persuade people; find the right people and work with them. If you are not working inside the core of AI, you typically don’t understand the industry’s demands and pain points. Don’t get fooled by the sexiness and fanciness of the AI industry. Money may talk in the short term, but technology speaks louder.
AI competition is real. When we told people this is a market with 1-2 players, people asked, “Is this a big market?”; when we listed the pros and cons of MatrAIx versus the other players (people like to call them competitors, but actually no, we are just players in the same domain, and our technical approaches and target users are quite different), they said this market is too crowded and noisy. But the market has just emerged: OpenAI becoming a successful company doesn’t make Anthropic, Perplexity, Mistral, or Alibaba Qwen fail. Instead, it builds a sustainable ecosystem where many players have different focuses. There might be some friction, but usually everyone’s goal is to bake a bigger cake instead of stealing a small slice from other people’s plates.
Want to join us?
I really need to get back to building (my team is calling me!). Starting as President and COO of MatrAIx means that I need to be very focused, ambitious, concentrated, money-driven, responsible, and to have high endurance for failures and rejections. I feel the responsibility as well as great pride in starting MatrAIx with my crew, who share the same goal.
We are vibe-hiring (which doesn’t mean we are unserious). If you feel the right vibe and want to join us, please reach out to yuexinghao@matraix.ai. We are creating a new paradigm for evaluation, so MatrAIx’s interviews should also follow this new paradigm. We pay less attention to LeetCode and beautiful resumes, and more attention to the ability to use coding agents and to communication skills (in human-agent prompt interactions and human-human verbal communication).
Cheers, YH.