Glasp’s note: This is Hatching Growth, a series of articles about how Glasp organically reached millions of users. In this series, we’ll highlight some that worked and some that didn’t, and the lessons we learned along the way.
If you want to reread or highlight this newsletter, save it to Glasp.
Popular Episodes at a Glance
#3: How We Rode the AI Wave with Side Projects Before It Exploded
#17: How We Grew ChatGPT Traffic 37x Using Our Own Server Logs
#18: How a Two-Person Team Uses AI to Run Multiple Products With 3.5 Million Current Installs
Hi, Kei and Kazuki here.
Glasp is two people. Between us we maintain browser extensions, a web app, native mobile apps, a separate AI product, and, since this June, a research practice that has put out twelve papers.
Last time we promised an issue on support, research, and the agent workflow behind them. We are splitting it. This one is research, because this week Kazuki is in Minneapolis presenting the first of our papers to pass peer review. Support and the agent workflow come next.
In this issue:
How it started. A widely shared number about our own growth, and why we turned it into a paper instead of another post.
What the data says about people. Language models agree with each other about what matters on a page. Readers do not agree with them.
How we pick a question, and the four tests a question has to pass.
What AI does in our research, and the four things it does not.
Why it was fast, and what “fast” leaves out.
What we got wrong. A headline that was a bug, a finished paper that was not new, and two more.
What the research changed in the product, mostly by stopping us from building things.
Questions we get, including whether it is ethical to use AI and user data this way.
The first part is open to everyone. The rest is for paid subscribers.
From our friends at APIdays
APIs run the world. APIdays brings the people behind them together.
Join developers, architects, and tech leaders shaping the API-first world at APIdays. Explore expert-led events covering API strategy, governance, security, AI, and the technologies transforming modern businesses.
How it started
Glasp did not set out to publish papers. It started with one experiment.
At any given time, Glasp has a lot of growth tactics running at once. At the end of 2025, answer engine optimization looked like one we could actually win at. More people were asking AI assistants what was inside a YouTube video instead of searching Google, which is exactly the kind of question our hundreds of thousands of YouTube Q&A pages were built to answer, and at that scale SEO had few levers left we could pull. So we bet on optimizing for ChatGPT, and we ran the experiment. (The full story of that bet is in Hatching Growth #17.)
It grew our ChatGPT referral traffic 37x, and he wrote it up as a long guest post for Sean Ellis. It reached far more people than we expected. People shared it and argued with it in the comments, and a few emailed to say they had built something by following it.
The post had a weakness we knew about. It made a good story, but it was not a controlled experiment, so it could not say how much of the growth our changes had actually caused. Kazuki suggested writing it up properly and putting it on arXiv. Kei had never written a paper and mostly said “sure.” Kazuki started anyway, and it became Disentangling Answer Engine Optimization from Platform Growth, which separates the effect of our changes from the growth of ChatGPT itself.
That is the first lesson of this whole issue, and it is free: the number about yourself that everyone is sharing is the one you should check hardest. A story can be true and still not tell you what caused it.
Why a two-person company does research
We had always half wanted Glasp to be a research company, a place that studies interesting questions and not only ships features. Part of that is plain curiosity. We run reading groups, and we read a lot every day, books and articles alike. Glasp is a product for exactly that. In what we read, the best discoveries usually came from someone following a question because they wanted to know the answer. And part of it is watching companies like OpenAI, where frontier research is the product.
There is also a practical reason. Innovation tends to happen when frontier research meets a real product. People like to talk about connecting the dots, and when you look at the dots behind most new products, one of them is almost always a new technology. We have seen this work once already. As we wrote in Hatching Growth #3 and #4, we were convinced well before ChatGPT that AI was about to take off, and we placed our bets early with AI side projects such as DALL·E-dle, a Wordle-style game where you guess the prompt behind four DALL·E 2 images. That is why we could ship a ChatGPT extension on the day ChatGPT launched, and then the first YouTube summary tool while most people were not yet paying attention to AI. That head start is most of why YouTube Summary grew. Knowing where attention is about to go is a research problem, so we treat it as one.
In practice, that also means Claude reads arXiv for us, often. When a paper on marketing, AI, or agents reports that something works on public data, we have it unpack the method and test it on our own YouTube content and pages. If the result holds for us, we may be one of the first teams outside the authors to use it.
And it fits why Glasp exists. Our mission is to leave knowledge for the people who come after us, which we wrote about in Hatching Growth #16. Studying something carefully and publishing what we found is about as direct a version of that as there is.
What the data says about people
Most studies of highlighting are small lab studies. We have years of highlights that people made while actually reading, for their own reasons. That means we never have to recruit anyone or collect anything new. The data already exists, and it describes something otherwise hard to observe: what people actually pay attention to.
Here is one result. Claims that language models make everyone think alike are usually tested against human judgments collected for the study, which makes the human side an artifact of the design. We used a human reference nobody built for the purpose: 2,523 sets of highlights that readers made across 120 web pages.
On a typical page, two readers who each marked 14 sentences shared about 4. Two models shared almost 9. Across 18 model configurations from 11 vendors, models agreed with each other far more than readers agreed with one another, and no model agreed with readers detectably more than another reader did (Language Models Agree With Each Other, Not With Readers).
If you use AI to decide what matters in a document, to summarize it, compress it, or pick the key passages, this is worth sitting with. A panel of different models is not a panel of different readers.
This is our second paid field note. The regular newsletter remains public, and we will keep publishing these when we have a concrete system and real implementation detail worth showing.
Below: the rest of what the data showed and the agent research that started at a hackathon, how we pick questions, exactly what we hand to AI and what we keep, why it went so fast and what that hides, four things we got wrong, what the research changed in the product, and our answers to the questions people ask most.





