I took a look at the preprint, Negligible participation in the scientific literature of AI unicorn startups, over the weekend and, whilst I think, who the hell am I to critique the work of John Ioannidis and team, I also think this is a weird preprint [1].
- The study population is AI-related unicorns in 2025, but this is really a mixed bag of companies. It includes frontier labs, traditional computer-vision companies, robotics companies, infrastructure providers, biotech companies, SaaS companies, legal products, fintech, food technology companies, etc. [2]
- Companies like Google DeepMind and Meta are missing from the list because they’re publicly owned.
- The authors identify 3,410 publications containing at least one startup-affiliated author, but remove 1,333 papers because they decide a startup has only made a sufficiently “leading” contribution when it has a first author, last author, or an author on an alphabetically ordered paper. I’m unclear how valid that approach is?
- The denominator is a bit weird. It compares 317 private companies with the entire global AI literature. This leads to great headlines, but I’m uncertain how meaningful it is beyond that, especially given how AI research has exploded in recent years and how thin much of the research is compared with, for example, a clinical drug trial.
- Whilst we’re clearly in bubble territory for most AI company valuations, I’m not sure scientific publishing is a particularly useful test of whether most of these companies are worth what investors think they are. For a biotech company, the underlying science is often the product, and you need evidence that it works. For Harvey, Canva or an AI infrastructure company, customers using and paying for the product is more relevant.
- For the frontier labs, transparency is an issue. But many frontier AI companies are developing a parallel technical literature outside conventional scholarly publishing, such as model cards, system cards, benchmark results, technical reports, research blogs and similar outputs.
- The top 10% of firms account for 96.8% of citations, but that’s mostly OpenAI (39.4%) and MEGVII (26.6%), a Chinese computer-vision firm from an earlier generation of AI with a fairly conventional academic publishing culture.
[1] See also Science write-up of preprint: https://www.science.org/content/article/ai-s-top-startups-are-barely-publishing-their-research
[2] As you would expect, the frontier-model firms are on the list, e.g. OpenAI, Anthropic, xAI, Safe Superintelligence, Mistral AI, Thinking Machines Lab, etc. But you also have infrastructure companies like Databricks, Cerebras Systems, Groq, SambaNova and Snorkel AI. There are defence, autonomous-vehicle and robotics companies such as Waymo, Cruise, Figure, Anduril, Shield AI, Nuro and Applied Intuition, where plenty of the research is going to be commercially sensitive. There are GenAI application companies like Harvey, Cognition, Anysphere, Sierra, Synthesia, Jasper and Mercor, plus OpenEvidence, Abridge, Hippocratic AI and Commure in healthcare. Other than the healthcare applications, it’s unsurprising none of these companies are publishing a lot of research. Similarly, I doubt many people would be expecting much academic AI research from Miro, Canva, Superhuman/Grammarly, Celonis, Vercel, Talkdesk or Automation Anywhere.
Also posted to Substack: https://substack.com/@pubtechradar/note/c-311938893



