Zhao Changpeng bets on third-year Chinese college students for $11 million as education agents

Author: Founder Park
Original title: Zhao Changpeng voted for a third-year Chinese college student, $11 million seed round to become an education agent
Third-year Chinese university students, $11 million seed round. Silicon Valley students are currently financing the highest product.
VideoTutor, an educational agent product aimed at K12's education agent product, announced today that it has completed a $11 million seed round with the main theme of generating personalized teaching/explanation videos in one sentence. This round of financing was led by YZi Labs, and Baidu Venture Capital, Jinqiu Fund, Amino Capital, BridgeOne Capital, and a number of well-known investors participated.
It is also the first AI product company YZi Labs has invested in.
Founder Kai Zhao (Kai Zhao) said that VideoTutor received approval and support from Changpeng Zhao and the YZi Labs investment team, and eventually YZi Labs led this round of financing. They got more than 10 TS (Letter of Intent to Invest) and ultimately chose these companies.
The first version of the product was launched on May 14 (launched at the Founder Park Product Market), which was approved by the market and verified by PMF. In less than 5 months, the $11 million seed round was completed.
According to Kai, the core reason they were able to get this financing is that, on the premise that the direction is right, the “Little Genius Team” used visual learning methods to solve the pain points of studying in the US college entrance examination at the K12 circuit.
“This field is more suitable for young people. In addition, they have very good hands-on engineering skills, and the founder himself has very good insight and experience, and the execution is very fast.”
Not just them, Cursor, Mercor, Pika, GPTzero, etc., Silicon Valley university students are using AI products that have set new highs in financing one by one to refresh everyone's perception of AI entrepreneurship.
Entrepreneurship in the AI era is really a bit different.
We talked to these young people at VideoTutor and wanted to know why they were able to get this seed funding round and what kind of changes are happening in Silicon Valley startups today. Also, why do they want to recruit employees from major domestic manufacturers so much.
Interviewees: CEO Kai Zhao, CTO James Zhan.
Interview & Editor | Manho
Here's the interview, edited by Founder Park.
01 K12 track, visual learning is the real direction
Founder Park: So many organizations are very optimistic about you. In your opinion, what is the core point that moved them?
Kai: First of all, I think the direction is right. The AI education circuit has great potential and prospects. The fields of education we are entering are SAT and AP in the US college entrance examination. The target user group is K12 high school students. The gap between us and this user group is very small; there is basically no generation gap. We have gone through the entire exam preparation study cycle, know where the pain points of exams and exam preparation are, and can make a product that really solves the pain points of this group of people.
Second, the team is excellent. James is from Gemini and is a core engineer in AI engineering and algorithms at Google. I myself have three educational entrepreneurship experiences. I started a business as an educational software business from the beginning of my freshman year, and participated in the creation of MathGptPro during my sophomore year, where the project was selected for the Miracle Innovation Forum, etc. Experience in successfully creating educational products.
Third, the core of our AI education field is animation engines, and we are the core developers of VideoTutor, the team that knows the core technology the most, and can render animation engines very accurately.
The team itself has very good marketing genes and knows how to spread the word.
VideoTutor is very much in line with an investment consensus of mainstream American VCs, called the “Little Genius Team”, which means that this field is more suitable for young people. In addition, they have very good hands-on engineering skills, and the founder himself has very good insight and experience, and the execution is very fast. I think this is a common reason why all investors can be optimistic.

VideoTutor Listed on NYSE at YZi Labs EASY Residency Demo Day
Founder Park: What core problem in the education industry is your product trying to solve?
Kai: The current learning products on the market can be classified into two categories: active learning products and passive learning products. Passive learning products, such as Byte's Gauth, Chegg, AnswerSAi, etc., cover the scenario we call “Homework Help” (Homework Help). The learning link is very short, and mainly students pay to do homework answers.
VideoTutor, on the other hand, covers active learning scenarios. We don't need to consider students' motivation to study, because they have to study and take tests, such as the SAT and AP for the US college entrance examination. In this scenario, there is a large demand for visualization of pain points. 80% of the content of the US college entrance examination involves knowledge such as functions and calculus that require complex image rendering. VideoTutor's animation engine solves this scenario very well.

Moreover, the customer unit price in this field is very high. On average, 2.6 million students in the US take the SAT test each year, and there is a huge demand for payment. Offline SAT courses are expensive, not per package, but by the hour. On average, they start at $150 per hour, with most costing $230. Many students and parents pay to study. However, VideoTutor is a great way to transition or even replace teacher training, because there is almost no difference between AI-generated video and teacher training content at this stage. In this way, students can have their own AI personalized exam preparation teacher at the lowest cost.
Founder Park: What made you decide to make this product?
Kai: Actually, before us, there was already a team at Stanford called Gatekeep Ai. They also wanted to do visual learning at the time. I was already aware of the influence of this direction at the time. When starting a business a few times ago, the educational products that everyone made were basically connected to GPT's API, similar to a ChatGPT Wrapper product. But we discovered that, based on text Q&A alone, this type of product has a ceiling. As can be seen, businesses such as Chegg and Gauth are declining, and most of the scenarios have been replaced by ChatGPT because students can solve many homework problems by paying $20 to use ChatGPT.
Products based on the API package and optimization level have reached the ceiling.
However, multi-modal visual generation has great prospects, because there are so many visual learning scenarios in the field of the US college entrance examination. Unfortunately, Gatekeep got off to a good start, but it didn't continue because it was launched a little early, the basic model programming ability wasn't mature at the time, and GPT-4 wasn't released yet. Coupled with the mathematical animation engine involving rendering and algorithms, they didn't overcome it. But our team has mastered all the core development of the animation engine, solved this problem, and made the video rendering very accurate.
02 PMF: Users are very willing to pay
Founder Park: After you launched the product at the time, you also reached cooperation with several schools. In your opinion, when or what feature made you think “I did the right thing with this product, found the right pain points” and felt like you had found PMF?
Kai: You can think of it in three dimensions.
First, from the level of revenue metrics, up to now, VideoTutor has received API requests from 1000 companies, including all well-known large educational institutions in the US, and even domestic institutions. Also, there are many schools that want to buy services. The intention of the C-side user is more direct. One parent of a student is also an investor. After experiencing the product, he gave the product to all his family and friends to try it out, and everyone was willing to pay. Then he didn't know where he got my phone and sent me a text to vote for us. C-side users are very willing to pay.
Second, from the user's demand level. Why is online one-on-one tutor education in the US so rigid? Because parents feel that one-on-one teaching works well, they are willing to pay this money. Now, multi-modal AI technology can personify one-on-one teaching results, and questions are answered. Moreover, video lessons recorded by online teachers in the US are actually no different from AI-generated videos. This is what I call “demand shifting”. Recorded and broadcast courses that students spend a lot of money on are no different from those generated by my AI, so why not use AI? The cost is lower, and the teaching effect is better.
We have received very positive feedback from many students, and many teachers are also willing to spread the product. The broadcast completion rate and usage time in the early stages were very good. The 200 torrent users we have now screened are all early accumulations.
The third point is the taste and sense of a product. When you keep doing it, from the progress of the education industry as a whole, to the core requirements for students and parents to the evolution of the product itself, the whole logic is closed. So looking at these three dimensions, you think PMF is enough. At its core, the willingness to pay is very, very strong.

A partnership was reached with FIZZ
Founder Park: Many users want to pay, and others actively contact you to invest.
Kai: Right. In the field of SAT and AP, the willingness to pay is already very strong. The unit price for customers in this field is as high as 100 to 200 US dollars. Offline classes are more expensive, and may cost 800 US dollars. There are 2.6 million students in the US who want to take the SAT, and 37% of them take the initiative to pay. This is a market with strong willingness and demand to pay. Our products can achieve a very good shift in demand.
Founder Park: On the SAT circuit, for candidates, a real teacher and an AI, will they trust AI?
Kai: AI now answers questions at the level of SAT and AP in the US college entrance examination, and it is generally less likely to make factual mistakes. In this case, why is it better than an offline tutor? One is cheap, and the other is that students can keep asking any questions. Don't worry if you ask stupid questions, teachers will have opinions or be impatient, and they can study anytime, anywhere 24 hours a day.
Moreover, this market can be moved. After completing the US market, we can also move to Canada, the UK's A-Level exams, etc., and the demand for payment is huge.
Founder Park: What are you considering paying now?
Kai: We have a monthly subscription, and we also pay for learning results. I think AI can now pay for results. We might launch a package, where you pay $799, and we guarantee that your child will get a perfect score on the SAT math test.
Founder Park: But when you pay for the exam results, doesn't it also depend on the student's personal motivation?
Kai: This probably won't work in the domestic college entrance examination because there are so many college entrance examination points, and there are thousands. However, there are only 62 SAT test sites for the US college entrance examination, 50 of which are regular test sites. Most students have no problems, and the remaining 12 test sites can basically be mastered. Unless there is a real problem with this student's level of logic, there is almost no such thing as not being able to learn. Moreover, the efficiency improvement effect of AI is very obvious.
In fact, many online tutors in the US also have this service. You pay the teacher 1,800 US dollars, and the teacher tutors the child. The success rate is basically 100% because the SAT test site is fixed. As long as the student's IQ level is normal, there is basically no problem. However, the college entrance examination didn't work; there was no way to advance in the college entrance examination in the short term. Moreover, domestic college entrance examinations need to close the score gap, and there will be problems, but there are no absolute problems in the US college entrance examination, because it is more about checking whether you have mastered knowledge points.
Pay-for-results is also a model already used by teaching assistants before, and this precondition is met.
Founder Park: So in your pricing, would model cost be a problem? Is it a high proportion?
Kai: The customer unit price in our field is very high. They all start at $69 a month. The model cost is now very cheap, so it's not a problem. The education industry is not like the coding field; everyone is charging prices because coding requires supporting a very long context.
03 For products aimed at high school students, the web is the most important
Founder Park: Remember the last time you said that your first prototype took almost just over two months. How were you considering the entire development cycle at the time, such as the division of labor and deciding which functions to do and which not?
Kai: Everyone on our team agreed that we need to iterate fast, because the faster we get feedback from early users.
After the first version was posted on Twitter, it caused a lot of buzz and brought in a large number of users. However, many of these users are programmers, investors, or technology enthusiasts; we can collectively call them “early adopters of technology.” At that stage, the feedback we received from them was scattered and of little value. It is still necessary to select the real core seed users from such a wide range of users, that is, high-quality high school students, and then obtain useful feedback through consultation.
The core feedback we received was that video rendering must be 100% accurate, which is a top priority that needs to be optimized. Whether the UI looks good, or whether it supports different TTS sound choices, these features have all been cut by us. Back to the core of the product: what we do is learn knowledge in science scenarios, so the accuracy of graphic rendering is the core.
Founder Park: What was the trade-off when it took to generate?
Kai: At that time, the highest peak time was about 6 minutes. The main consideration at the time was that the explanation of common topics and explanations of knowledge points should not take more than 6 minutes. However, in subsequent feedback, we found that some students were not that good at learning, and wanted the content to be slower and more in-depth. We are aware that the length of time should not be limited; it depends more on the user's ability to learn.
Founder Park: How long can it last now?
Kai: The longest should be less than an hour, so you can keep breaking the casserole until the end. It is generated in real time while communicating, but this feature was recently introduced, and the first version didn't have it.
Founder Park: Is there a feature you wanted to do at the time, but then later found that it wasn't that important, so don't do it first?
Kai: For example, apps. At the time, I thought it was necessary to develop apps quickly, but later I discovered that most students in the US basically use laptops or iPads to study. Most K12 schools in the US give students a Chromebook computer. Computers are highly popular, and they also complete their homework on computers. High school students basically own a computer, and mobile phones account for less than 5% of the study scene, which is very low.
Founder Park: So if it's a product that focuses on education or student groups, the web app is the first thing to do, but the app isn't that important.
Kai: Yes, I actually already knew this data at the time. After all, I spent many years studying in the US. Later, we surveyed 100 students from tens of thousands of early users. More than 90 of these 100 students had computers, so we were even more convinced of this.
Founder Park: When you launched the first version, were you also targeting the K12 community?
Kai: Yes, they also targeted this group later. We don't compete with Gauth; we do more of an exam training scenario. A large number of high school students in the US themselves choose offline training or online learning platforms, and VideoTutor did a good job of shifting this demand.
Founder Park: Will K12 be your core user group for at least a year?
Kai: It should be a core indicator within two years.
04 Use big models, but don't just rely on big models
Founder Park: Can you give us a brief introduction to your current technical implementation plan? VideoTutor does a much better job of generating courses and diagrams than other video generation models, and even when many models can't even generate text accurately, your technology is amazing.
James: The videos we generate have both text and graphics. The approximate production process is: let the big language model generate text and corresponding animation instructions, then the animation instructions are rendered by our animation engine, and finally presented on video.
The text part is relatively simple. We let the big language model generate the text and then render it directly. But the animation part was generated by our own mathematical animation rendering engine. Its advantage is that it renders content such as axes and geometric figures very accurately, and this is where our core technology lies.
The current big language model only outputs text. The set of agents we made is equivalent to giving the big language model a sheet of paper and a pen, so that it can draw the appropriate teaching animation it can imagine. The part I drew was all our technology.
Founder Park: How was the final synthesis of the entire video, including audio and video, handled?
James: At first, the user will receive a prompt, such as “What is Pythagorean theorem?” In the first step, we let the big language model infer all the scenarios. Generally, we will specify 3 to 5 scenarios, depending on the difficulty of the problem. The model then generates a rough script for each scene. Next, a second inference is made based on the script for each scene to generate the text of the scene, corresponding patterns, and vocals. Vocal text is then synthesized using TTS.
Finally, we put all the scenes together to form a complete video.
Founder Park: I understand that the first edition was a plan like this. Now that we've added a ready-to-interact process, has the generation process changed too?
James: There have definitely been changes. Now, in order for users to see the content as quickly as possible, we will first generate the first scene so that the user can watch it first, and the later scenes will continue to be rendered in the background. When a user asks a question, we convert his vocals into text, then hand this text along with the content of all previous scenes to the big language model for reasoning, and let it plan the next teaching scenario. The rendering process for subsequent scenes is the same as before.
Founder Park: If a user hears a question for a minute, he'll ask it directly. Once you have received the questions, return the user's questions to the model for processing along with what was mentioned before. In this process, after users have asked questions, will the animation continue to broadcast or will it stop?
James: Our delay has now been reduced from 20 to 30 seconds at the beginning to less than 5 seconds. In terms of interaction, we'll make some transitions so that users don't pay too much attention to these 5 seconds, and the whole process will be more seamless. Within 4-5 seconds, he was able to see a new presentation based on his question.
The design at this stage is that the AI teacher will say, “Uh, I'll consider it,” and then wipe the blackboard, just like a real simulated teacher. If you think there's a problem with the explanation, then I'll erase it and write it again for you. This process will feel more natural.
And we don't just passively wait for users to ask questions; we also do quizzes along the way. We make inferences based on quiz feedback and user questions. Moreover, we are not completely free to use the microphone; instead, we need the user to actively turn on the microphone and turn it on and off.
Founder Park: So based on this mechanism, you can generate an explanation of up to an hour.
James: There's no limit, to be exact; if he keeps having questions, he can keep asking them.
Kai: Yes, there are no pre-set limits. In fact, VideoTutor is moving in this direction, and with the advancement of multi-modal AI, we are not creating needs, but are better meeting existing needs. Look at offline real-life education, why are American parents willing to pay a lot of money? Because the US education and training industry is more about one-on-one teaching, it starts at $100 per hour. It's because offline teachers can ask guided questions, so I can observe where you aren't doing it, and then ask you. VideoTutor also tries to achieve the teaching effect of a real teacher, so that every child can interact and teach in real time.
Founder Park: Will students be asked to turn on their cameras in class?
Kai: Not really. Whether students turn on cameras depends largely on US privacy laws. The product doesn't often design features that are forced to be enabled; whether to enable it depends on the student's wishes. The main interaction was also through questions and voice feedback.
Founder Park: Technically, are you using a strategy that combines the small model with the big model in the cloud, or what?
Kai: It's a collaboration. We have an internal data set and now have over 100,000 pieces of video data. The best of these data are all manually labeled twice and then used to train fine-tune the model. For example, we have over 8,000 SAT sample training data. These small, fine-tuned models will work with general-purpose commercial models like Claude and Gemini in the cloud.
Founder Park: Will using Claude, Gemini, or GPT affect the core performance of the product?
Kai: We're mainly involved in the K12 field, and the level of the basic model is good enough. However, in order to ensure 100% accuracy, we will call the two models to check at the same time. If the answers of the two models match, then there is basically no error. In terms of code generation, Claude is still the main one, and its coding ability is relatively good.
Founder Park: Where are the product's technical bottlenecks right now? Is it model ability or code generation?
Kai: Model ability is one part of it. There's also rendering, which is now under 5 seconds, and can be even faster as more GPUs are deployed. The other is the ability to remember for a long time. We need to accumulate data on students' learning behavior over a long period of time. We know what knowledge points this student doesn't understand. For example, if you forget the knowledge points you learned a month ago, you can remind him again.
James: We actually put a lot of effort into rendering time and have been making technical breakthroughs, from the first 2 minutes to 1 minute to less than 10 seconds now. Our ultimate goal is to be able to render with almost no delay. As soon as the user asks, the results come out as soon as the inference is over. This is a challenge our team is currently overcoming, but we have found a new direction.
05 Don't watch the broadcast rate, just watch the final exam score
Founder Park: How do you measure the product's core metrics at this stage? How can you tell if a video is useful to users?
Kai: One of the core metrics is the test. In the new version, after watching the video, there will be a quiz at the end. If you do it right, it will prove you understood; if you didn't do it right, it would prove that you didn't understand.
There is no way to just watch the broadcast rate for learning results; some students may understand it by reading half of it. Give him a test while he is half watching, and if it passes, he doesn't need to read the rest. The core indicator of our product is to see how many students improve their scores here.
Founder Park: But he finished his final exam in a different scenario. How did you get this result if he passed?
Kai: When it comes to American product culture, users get good results after using the product, they will spontaneously share it. Many students take the initiative to share their experience and results after completing the SAT with VideoTutor. We will also make them campus ambassadors for secondary dissemination.
We have 20 campus ambassadors made up of high school students. In fact, you see Mercor was very successful in the early days, using a typical “user success story” model. Mercor helped many Indian programmers find jobs in the US in the early days, and then they would contact these users to shoot a user story and tell them how they used Mercor to find jobs. This has formed a great spread of word of mouth. VideoTutor also makes sense. What we want is for more students to achieve very good results after using the product, and then share these students' experiences as user stories.
Founder Park: What are the main channels for students to share?
Kai: Students are mostly on TikTok, parents are in Facebook groups.
Founder Park: If you put the time in terms of half a year or a year, how do you plan to grow your product?
Kai: I think essentially, VideoTutor's core is still a C-end user product, and word of mouth is very important. Many successful AI applications in the early days depended on the word of mouth of seed users. For example, designers felt good about using it and spread. For us, the core indicator is how many SAT candidates got a high score after using this product, and then spread it to other kids and parents. Parents mainly use Facebook and Instagram, students use TikTok, and we broadcast on these platforms. When this kind of consensual word of mouth is formed, it is natural for school teachers to recognize it. We were known by so many schools in the early days because many teachers thought it was good when they used it and recommended it to the school's procurement manager. Therefore, the core is still word-of-mouth communication among C-end users. How many children improved their scores after using it is a key indicator.
Founder Park: What is the approximate status of the new version and what is the timeline for launch?
Kai: We want an official public release within two months as soon as possible. At that time, students can answer questions with very low latency, and graphics rendering of science scenes can be 100% accurate. Of course, we won't be covering competition scenarios or complex university knowledge such as linear algebra for the time being, and more in the K12 field.
Founder Park: VideoTutor What's the current barrier or moat?
Kai: I think there are a few things. The first is the data flywheel. Behind the video is all code. Good video data generated by users can be re-trained to fine-tune the model after being labeled twice. The more data, the better the video. In addition, there is learning behavior data. We know which knowledge points of different students are weak, and we can set up a data flywheel. The more people use it, the more students understand the product. The second is leading technical advantages, such as animation engine algorithms. Although the algorithm itself is not the core advantage, as we iterate rapidly and there is more data, the advantages will become more obvious.
Third is the brand. VideoTutor has become a leading brand in the field of AI education in the North American parent community, and parents' trust is also an invisible barrier.
Founder Park: What kind of product do you expect VideoTutor to eventually grow into in three to five years?
Kai: In the future, we want VideoTutor to be an AI teacher for everyone to learn science knowledge. We only do science. I think in the future it will surpass many of its neighbors. Duolingo is a world-class language learning product, but in the STEM science scene, world-class products have never appeared in the past because science requires too much graphic rendering. Now that the technology for the basic model is ready, I think the next “Multiple Neighbors” will be born in the science scene.
06 Recruit people, especially those who want to come out of big domestic manufacturers
Founder Park: How many times have you started a business before, what do you probably do?
Kai: I'm in my junior year. When I was in my freshman year, I started a business with James to make educational products, and received 200,000 US dollars in angel investment. Despite that failure, I learned a valuable lesson: you can't fall into homogenous competition. The apps we made at the time had many similar products on the market, so they had to fall into competition in the early days, making it difficult to charge fees.
The second time I started my business, I joined another team, MathGptPro, as a co-founder, and stayed for a few months. At that stage, I learned how to read product metrics, how to build products, and how to expand users. It was also at that time that I came to the conclusion that text-based answering educational products had come to an end. Because it's no different from ChatGPT, and in the past, structured knowledge question banks that used to be done with homework help at a great cost have also been replaced by the editing ability of big models. So when I started a business for the third time, I knew that visualization is an inevitable trend.

Zhao Kai's group photo with Sam Altman pitch at Harvard University
Founder Park: Apart from making you aware of the limitations of text-based products, did the past two experiences help you as a VideoTutor now, in terms of the team or otherwise?
Kai: It helped a lot.
The first point is to better determine whether the direction and product have a future. I will judge the evolution direction of the entire product by looking at competitors' website traffic and revenue.
Second, in terms of product construction, it is possible to better judge the pace of product development, including product design, front-end and back-end docking, and which indicators to look at.
Third, team management and organizational culture skills. I have established a more complete management system, including division of labor, rewards, and option distribution for each student. Also, I learned how to finance. We completed this round of $10 million financing in less than 20 days.
Founder Park: How many people are on your team now?
Kai: 6 people, everyone lives together.
Founder Park: How was the team originally built?
Kai: James and I have already started a business twice. We both graduated from the same school and built an app together in our freshman year. In my sophomore year, I started a business with two other people, and they all knew each other. When we realized that this technology could bring about a very big product vision, we teamed up to make this product. Everyone used to be alumni, including Nick, another partner on the team, who was also my college roommate.
Founder Park: You're also preparing to expand your recruitment now. What kind of people do you want to recruit?
Kai: We mainly recruit back-end, front-end, big language models, and UI/UX, and I hope they have experience. Because we have now gone through the trial and error phase and entered the stage of rapid product building, we need experienced people to help us grow.
Founder Park: We need experienced engineers, product managers, and growth leaders to take products from 1 to 10, or even from 10 to 100.
Kai: Yes, this is the stage. We expect to expand the team to 9 to 10 people, and the core priority is to recruit engineers.
The recruitment this time is likely to be domestic, so it's a mix of in-person and remote.
Founder Park: What would you like this person to look like?
Kai: We would rather have experienced it in some big companies, such as Byte and Meituan. Because Byte is a high-speed, comparative organizational culture that values young people. People who have been trained in Byte have good methodologies and abilities. After joining us, they can bring in these successful experiences and conduct integrated learning.
People who want to fight tough battles in major domestic manufacturers and have rapid iteration experience. We have already gone through the student entrepreneurship stage. We don't need to recruit newbies. We need to recruit more experienced people, but we are not the kind of complete “industry veterans.” Because veterans in the industry may have to take family into account, there is no way to do that. Therefore, those at the middle level, those who are young and able to roll are better.
We are willing to give excellent talents rich options. Although we have raised 11 million US dollars, why haven't we recruited engineers in the US? That's because we think domestic product strength and engineering capabilities are really good. This wave of 100% Chinese-run teams will create great products and compete internationally. Many AI application levels are now created by Chinese people, and domestic engineering capabilities are really impressive. This is also our advantage; we need to take advantage of the advantages between China and the US.
VideoTutor current detailed recruitment requirements: https://videotutor.io/
07 Silicon Valley college students have all started businesses with AI
Founder Park: Now, especially in Silicon Valley, the trend of college students starting businesses is particularly obvious. What kind of state are you seeing?
Kai: If you look at a fact, let's say the company with this round of 10 billion dollar valuation: Mercor, which focuses on AI recruitment, has completed new financing of more than 300 million US dollars, and the valuation is already 10 billion US dollars; Cursor is already a definite 10 billion US dollar valuation. Corresponding ones include GPTZero, Pika, etc. These are startup projects for college students; in particular, the founders of Cursor and Mercor are all junior college dropouts.
One characteristic of this wave of young people starting businesses is that competition is highly differentiated. They focus on doing things in a very narrow field; they don't do generic stuff. Mercor, for example, did AI recruitment, and initially only recruited Indian programmers.
The second point is the environment. The entire Silicon Valley capital environment and underlying innovation, such as Stanford, YC, and Peter Thiel's funds, all support college students to start businesses at the earliest stages. Whether you have mature ideas or not, they are willing to support you and provide a strong network of contacts.
Third, I think it's the quality of these college students. Both us and these college students from Silicon Valley have a very brave sense of adventure and a strong ability to learn. Many domestic students probably don't have this kind of brave and aggressive spirit. Because in Silicon Valley, you are inspired by the success stories of your peers around you, and the capital environment is willing to trust young people.
For me, I also compared costs and benefits at the time. If I choose to finish college and then find a job, I may not be able to pay for the costs of studying abroad at home, nor will I necessarily have a great return on my earnings. But if I choose to start a business, I can study like crazy when I'm young, and my life's possibilities are limitless. I wanted to start a great company since I was a kid.
Founder Park: Why can today's generation of college students start a 10 billion dollar company, while it was amazing that they might have sold a $120 million before? Is there an AI boom and bubble factor in this?
Kai: I don't think it's completely a bubble. Cursor has real revenue of 450 million dollars, which is very reliable. Behind this is the critical methodology and cognitive insight of this generation of young teams. Look at these teams, they all have excellent backgrounds, and they have very good learning abilities.
In the early days, Cursor relied on college programmers around them. These people were highly receptive to AI and gave strong feedback. The founder himself is also a little genius engineer. He has a deep understanding of users, and has strong engineering iteration ability. Four people started the product in the early stages. After they iterated on the product, they formed a reputation among users. With revenue, investors were afraid of missing out on the next Mark Zuckerberg, so capital helped again.
The bottom line is that many of the technologies in this wave of AI are new. Young people learn fast, are pragmatic, reliable, and dare to work, so they have the ultimate user understanding and ultra-fast iteration speed to beat traditional products. For example, before Cursor, GitHub Copilot did a great job, but why hasn't it been done? It's because of the user experience and speed of execution.
Founder Park: Is it fair to say that because AI is a new technology, many product perceptions also need to be viewed from a new perspective?
Kai: Yes, the younger generation has deeper perceptions and opinions than the previous generation of entrepreneurs, and can be closer to users. Today, mainstream AI users are all post-00s, and their learning and feedback are iterative and inclusive faster than the previous generation of entrepreneurs.
Therefore, cognitive iteration speed is central. In the mobile internet era, technology iteration is measured in years or quarters, but in the AI era, technology iteration may be measured in days. As a founder, you have to learn fast, and young people stay up late and work harder.
Founder Park: Earlier, some media said that many founders in Silicon Valley have also started 996. What do you think?
Kai: Some of my friends, who are white entrepreneurs, have raised a lot of money, also 996. Like us, they rent a big house and all live and work together. I think 996 is more forced by the environment. Now Silicon Valley is a bit like a gold rush. No one wants to lag behind, then they can only iterate faster than the product; they have to stay up late and iterate quickly. It's an environmental modification that forces people to do it.
Founder Park: Are these college students in Silicon Valley starting businesses any trends in choosing a racetrack?
Kai: I think whether we're educators or others, we all have a tendency to start businesses within our comfort zone. A comfort zone means that you know enough about the field and users. The founder of Cursor knows coding very well, and we are educating because we know this group well. Today's young people are starting businesses more within their existing cognitive comfort zone, and are no longer rashly jumping into a field they don't understand. Because this way you get feedback from users fast enough and accurate enough.
There is also cognitive superposition. We have been educating three times, and my perceptions are constantly being superimposed. This group of college students don't rashly do things they haven't done in the past; they all think about how to do better. They have a new generation mindset, are constantly iterating in their own cognitive circle, and are brave in creating opportunities.
Another point is that they have a brave and aggressive spirit. They don't deny themselves because of other people's denial. They have an “I don't care what you think about me” attitude, and they are very confident. Behind it is a culture of “high-speed experimentation”. I know my product isn't ready yet, but I don't care. Fast launch, quick iteration, and quick feedback.
Founder Park: When did this trend probably start?
Kai: I think it was a consensual success. When everyone saw that a project like GPTZero grew from a dorm room, iterated continuously, and then received capital support and user recognition, there were many successful cases of rapid trial and error, and rapid explosion, and a consensus was reached.
In a word, “better done than perfect”, completion is more important than perfection. Also, no one is too worried about competition. Many founders in Silicon Valley are willing to share their product ideas. They're not afraid you can copy them; I just need to iterate quickly. I think this wave of young people also has a good ability to tell stories. This kind of storytelling is not false, but is based on pragmatism and realism, plus their own vision for the future.
Founder Park: Market yourself first.
Kai: Right. I think the underlying idea is a sense of adventure and extreme confidence. Driven by this, they continue to be brave in trial and error, and are not afraid to say the wrong thing. Boldly talk about your product concept, boldly implement it, and change it if you make a mistake. This culture of not being afraid of trial and error has contributed to this wave of college students' entrepreneurship boom and success.
VCs in the US also look at college student projects, and YC regularly invests in college student projects every year.
08 Financing is the last thing VideoTutor has to worry about right now
Founder Park: If you go back to when you first started as a VideoTutor, what advice would you give yourself? What could be done better?
Kai: I think the pace should be a little faster. There's also team composition. The VideoTutor team went through many rounds of grinding. If I knew earlier, I would have formed a better team sooner based on the skill profile required by the product. I think that when starting a business comes back to the end, organizational ability is critical. I will spend more time on organizational skills: selecting people, knowing people, and using good people.
Today's team is suitable for growing from 0 to 1, but to make VideoTutors bigger, people with more work experience are needed to join them, bring their excellent experience and abilities to the team, and help the whole team grow together.
Founder Park: What kind of product or technical challenges do you think VideoTutor might face in the next six months?
Kai: I think one is rendering. To get to really zero latency, we need a breakthrough in engineering. The second point is growth. I think it's the taste of the product. There are many things behind this, such as whether the UI and interaction design are smooth and perfect, whether the functional interaction is bug-free, whether the visual layout is beautiful, etc. These are all tests for us.
James: I think in the beginning we positioned VideoTutor as visual teaching guidance for all subjects, but then we did it very vertically, just in the field of mathematics, because that's what we do best. Our math rendering engine is the most professional. The next key to break through is probably horizontal expansion. For example, how to bring the benefits of visualization to liberal arts scenarios? For example, explain “On the afternoon of Harvest Day, sweat drips down the soil.” This is the next technical point we need to consider.
Founder Park: Will subsequent expansion be bothered by the founder's background?
Kai: Not really. In fact, many big VCs have come to us. Like a16z, they don't take action too early, but only help when the team already has signs of success, so they know that the investment won't fail. We have excellent relationships with many big VCs.
Financing is the last thing VideoTutor needs to worry about, and the most important thing to worry about is the user ecosystem and product.
Twitter:https://twitter.com/BitpushNewsCN
Compare the TG exchange group:https://t.me/BitPushCommunity
Compare TG subscriptions:https://t.me/bitpush



