This is What Savvy AI Investors Look For
A wave of funding for generative AI is surging through the venture capital markets, but what makes an AI led company scalable and defensible?
LEARN
This is What Savvy AI Investors Look For
Published 14th Nov 2022 by James Taylor & Caitlin McCartney
Make your website intuitive
Discover the Particular Audience platform.
Learn more
In 2017, Andrew Ng (Industry AI leader, Adjunct Stanford Professor, founder of Google Brain and former Chief Scientist at Baidu) said, “Artificial Intelligence is the new electricity. Just as electricity transformed almost everything 100 years ago, today I actually have a hard time thinking of an industry that I don’t think AI will transform in the next several years.”
The technological and financial opportunity AI presents should have been clear to investors - according to IDC, worldwide spending on AI-Centric Systems Will Pass $300 Billion by 2026. With its seemingly limitless applicability, many fields are rapidly embracing Machine Learning and AI, from education, to finance, to healthcare, to ecommerce and so much more.
Oft quoted English science-fiction writer, futurist and inventor, Arthur C. Clarke (most recently by Packy Mckomick of Not Boring) wrote that "any sufficiently advanced technology is indistinguishable from magic". One of the best examples of magic tech in recent memory is undoubtedly generative AI - the ability to create images and articles from short text prompts was the stuff of science fiction a few years ago.
With the release of GPT-3 and Dall-E 2 by Open AI and a flood of exciting rival tech, a sudden wave of optimism reminiscent of 2021 is surging through the VC capital markets as investors pour funding into generative AI. Recent news of multi-million dollar raises for companies such as Jasper and Stability AI reflect the dramatic way this technology has captured the imagination and excitement of not only the general public, but also venture capitalists in the otherwise dreary markets of 2022.
General Partner at Unusual Ventures Sandhya Hegde writes that we are now witnessing “the third wave of Applied AI in SaaS: generative software…scaling unique human-like work output across modalities, including text, image, voice, code, music, 3D models."
While access to generative AI via open APIs is exciting and has already spurred many new ventures, companies built on open APIs accessing models trained on publicly available data (image and text especially) will struggle to remain differentiated.
If data is publicly available then anyone can access it, and over time software models tend to become commoditized as new entrants join the market, meaning that generative AI companies pose significant mid-term risks to investors.
So what are savvy investors looking for in an AI investment in 2023?
Data IP and true data network effects.
What does that look like? Differentiated, proprietary data sets and unique access to scalable, real-time data allows a company to build a long-term competitive moat around its technology.
That’s a lot to break down.
Before we do that, let’s quickly explain AI, Machine Learning, and the role of data.
What is AI and Machine Learning?
AI is the general field of intelligent-seeming algorithms, it involves programming computers to make decisions for themselves, mimicking or simulating human thought or behavior.
A large contingent of AI research today is focused around language and image readability, and generative AI is (understandably) creating a lot of hype.
Machine Learning is a subset of AI that deals with the creation of computer programs that can learn and improve on their own. A class of data-driven algorithms enable software applications to become highly accurate in predicting outcomes without any need for explicit programming.
If you’ve shopped on Amazon or watched something on Netflix, those personalized (product or movie) recommendations are great examples of machine learning in action.
What’s the role of data in AI and Machine Learning?
Data is used to train models in AI and Machine Learning. It’s also used to test and validate these models.
Computer scientist and Co-founder of Y Combinator Paul Graham recently tweeted:
This is something Andrew Ng (more recently founder of Landing AI) also feels passionately about. His campaign for data-centric AI aims to shift the focus of AI practitioners from model/algorithm development to the quality and scale of data they use to train models. Andrew believes that data is central to the future of AI and Machine Learning - not just model development. In fact, models in many applications are a solved problem, with most having been written as long ago as the 70s.
Making ML models more performant in real time computational environments requires bigger, cleaner, better and stronger datasets - but if those data sets are public and available, the models, while performant, will be commoditized.
In order to create genuine differentiation, a company must further develop their models through proprietary data.
What is proprietary data, and why is it important?
Proprietary data is data that is owned by a specific company or individual. This data may be confidential and may not be shared with others without permission from the owner.
- It allows companies to train their algorithms on a more diverse dataset, which in turn leads to more unique results. - It gives companies a competitive advantage over other companies that do not have access to the same data. - It allows companies to create customized models that are specific to their own business needs.
Partner at Canaan and respected deep tech investor Rayfe Gaspar-Asaoka believes that:
“To deliver on the promise of disruptive change, AI must create differentiation… There is a strong feedback loop between AI algorithms and data. The better the data, the better the algorithm will perform at future predictions. And the better those predictions, the better the output data, which is then fed back into the algorithm…a company with even the slightest head start with a better proprietary data set or algorithm will have an ever-increasing advantage over their competitors. This winner-take-all characteristic of AI is one of the things that makes these companies so powerful.”
Martin Casado of A16z famously wrote about the (empty) promise of data moats - it can be easy for a company, platform or product to generate data and claim they have a moat of some form, however this is often incorrect. He concludes that for data IP to be defensible and scalable, it has to be scant, unique, rare, high quality and difficult for others to obtain - ergo ~ideally, proprietary.
The Web is not proprietary, is it?
Generative AI in its current form trains on large bodies of publicly available data, and as models become a solved problem, they risk becoming commoditized.
Unique data becomes the defining factor for the future of a platform or product. This is nothing new.
We know from other domains, especially within the walled gardens of big tech, that collaborative (wisdom of the crowd) data used to help link entities is often what makes those platforms so uniquely intuitive, engaging and sticky. Over and above what they might achieve by only understanding the relationship and shared contexts of textual and image metadata.