Is My AI Startup Just a Wrapper? The Test Founders Actually Use

Is My AI Startup Just a Wrapper? The Test Founders Actually Use

August 10, 2026


TL;DR: You're a wrapper if the model does the work and you do the interface. Based on 200+ founder interviews on the PMF Show, the founders building durable AI companies pass four tests: they own context the model can't retrieve on its own, they operate where accuracy is legally enforced, they sit outside the roadmap of the foundation labs, and they carry the consequences when the AI is wrong. Run all four on your product this week — passing three is not enough.

After interviewing 200+ founders on the PMF Show, the wrapper question has become the single most common source of founder anxiety in AI. It's a fair anxiety: a large share of AI startups genuinely are a prompt, a UI, and a markup, and those companies get compressed the moment the underlying model improves or the labs ship the feature natively. But the founders on the show who are clearly not wrappers can articulate exactly why in a sentence or two. This article covers the tests they use — including the ones that disqualify ideas before a line of code gets written.

What actually makes an AI startup a wrapper?

A wrapper is a company where the intelligence lives entirely in someone else's model and the company contributes packaging. The tell is substitutability: if a competent team could rebuild your core value with the same API in a weekend, the model is the product and you're the skin.

Dan Mishin, founder of Manifest, used the term as an explicit disqualifier when choosing what to build. He'd decided he wanted meaningful, hard work, and he ruled out two categories by name:

"I knew I didn't want to build enterprise SaaS. I knew I didn't want to go and build a GPT wrapper that does something fun. I wanted to solve a meaty, difficult problem in a regulated industry, and legal just perfectly fit those criteria." — Dan Mishin, Manifest

What's useful is that "wrapper" for Mishin isn't a technical category — it's a difficulty category. His filter was meatiness. He went on to raise a $60 million Series A building AI-native legal practices, and his reasoning about why legal qualified is worth sitting with:

"Ask a healthy person what they want in life, they tell you a thousand things. Ask a sick person what they want in life, and they say one thing, right? Same applies to legal situations." — Dan Mishin, Manifest

Acute pain in a regulated domain is structurally hostile to thin products. Nobody buys a fun GPT wrapper for their green card application.

Key stat: Manifest raised a $60M Series A on a deliberate anti-wrapper thesis: hard problems, regulated industry, acute pain.

Test one: do you own context the model can't get itself?

This is the most rigorous version of the question, and Cos Nicolaescu of Accrual gave the clearest formulation of it on the show. He starts from why AI coding tools work so unusually well:

"the reason why it works so well is exactly that, which is I have all the context in the world in my repository. That is my world. I don't need anything else because that's how code works and I'm able to iterate through, structure things, test them, have that feedback cycle be as close as possible, and know the entire universe." — Cos Nicolaescu, Accrual

Coding agents aren't magic; they operate in a domain where the complete context is sitting in one place, machine-readable, with a fast feedback loop. Most industries are the opposite. And that's where wrappers get exposed:

"I think that is very different where I have an agent that extracts some information from a document and then gives you the result for that, and then another process that takes those, and wants to use it. But how many of those documents were related? And how much do you know about your client as a result of that? And what do you do with nuanced situations if you just get some raw numbers that are OCRed? You just lose all that context and all of a sudden, the capability of the model is diminished significantly from what it could do." — Cos Nicolaescu, Accrual

That paragraph is the wrapper test stated precisely. A wrapper does one extraction and hands back a result. A real product assembles the context that makes the extraction meaningful — which documents relate to which, what's true about this client, what the nuanced cases require.

"In this world, I think having that context is not only useful, but critical. I think without it, you will not get the gains of the promise of AI that you see in other industries without doing that." — Cos Nicolaescu, Accrual

The practical question for a founder: if you removed your product and gave a customer direct access to the raw model, how much worse would their outcome be? If the answer is "somewhat less convenient," you're a wrapper. If the answer is "they'd be missing the context required to get a correct answer at all," you're not.

Test two: are you building where the labs are already headed?

Nicolaescu also walked through his idea maze out loud, and the eliminations are more instructive than the destination. First, the foundation layer:

"We shouldn't do anything that foundational labs are doing. We're not research people." — Cos Nicolaescu, Accrual

Then the layer most AI startups actually occupy — and this is the one worth reading twice:

"Next level was things that are adjacent to the models that the model companies will likely end up doing, even if it takes them slightly longer. Probably not a good business decision, even if it's technically interesting, and you can do that." — Cos Nicolaescu, Accrual

That phrase — "technically interesting, and you can do that" — is an accurate description of most wrapper companies. They're buildable and they work. They're also sitting on a roadmap someone else controls. Nicolaescu ruled out consumer next — different growth motion, not his skill set — and landed on what he and his co-founder were genuinely differentiated at:

"The things that we're very good at are gnarly problems with complex workflows, lots of moving pieces that all are intertwined and have to click well together, super high reliability, super high accuracy, super high security." — Cos Nicolaescu, Accrual

That list is more or less an inventory of things a general-purpose model won't do for you. Gnarly, intertwined, high-stakes workflows are where wrappers can't survive.

Test three: who is liable when the AI is wrong?

The sharpest structural answer came from Rafael Broshi of Notch. Notch started five years ago as an actual insurance business — a managing general agent selling policies — and built internal software to service its own policyholders. Then it flipped: closed the insurance product, kept the software.

That origin produced something a wrapper can't buy:

"It was always more than a GPT wrapper that helps you answer FAQ. Because we were the one at stake here. It's not like you're selling a product and then you say, hey, good luck. I gave you the tool. You're going to be fine with it or not. It's that we were in charge of our own destiny. If someone would have sued us because our AI made a mistake, I'm the one who was supposed to handle that kind of lawsuit." — Rafael Broshi, Notch

Liability is a design constraint that changes architecture, not just marketing copy. A team that eats its own errors builds verification, escalation, and audit trails into the core. A team shipping a wrapper adds a disclaimer.

This is the test most founders can answer honestly in one sentence. When your product produces a wrong output, who pays? If the answer is "the customer, and they knew the risks," the market will eventually price you as a wrapper regardless of your architecture.

Key stat: Notch built its product while personally on the hook for its AI's mistakes as a licensed insurance provider — then productized it.

Test four: is the market actually opening, or are you educating it?

Being a wrapper is sometimes a timing problem rather than a technical one. Broshi's earlier company sold insurance for social media accounts and crypto assets, and the lesson he drew is unambiguous:

"what you learn from that experience is that it's very, very, very hard. And I don't think it is a startup's job to change a market or create awareness, OK." — Rafael Broshi, Notch

His read on why the current AI moment is different is that ChatGPT did the market education for everyone, and in doing so reset the board:

"That made everyone open their eyes and say, wait, I can now offer better support. I can now cut costs. I can do things that everyone tried to do for fifteen years with NLP, et cetera, but didn't work. And that, also meant that every company that has tried to do that up until this point and have poured millions of dollars or hundreds of millions of dollars. Now at that moment, there was an even playing field." — Rafael Broshi, Notch

Ali Khokhar of Amigo AI made the same call in real time. A month after ChatGPT launched he started mapping which labor categories would fall:

Never miss a founder's PMF story

Subscribe to The PMF Show
"each of these categories is going to get picked off one by one over the next five year window, let's say. And so it was very, very clear and obvious, at least to me at the time. Because it's what? November '22, right? ChatGPT comes out. February '23, I quit my job and I was like, I'm going to go build a company in this space." — Ali Khokhar, Amigo AI

Key stat: Khokhar quit his job in February 2023, three months after ChatGPT's launch, on a five-year category-by-category thesis.

The wrapper implication is subtle. A level playing field is an opportunity and a warning: if incumbents' spending advantage was wiped out, so was yours. Everyone got the same models on the same day. Whatever you build on top has to come from somewhere other than model access.

What if you fail the tests — how do you tell early?

You force a testable hypothesis, and you do it cheaply. Astro Teller, who runs X, Alphabet's moonshot factory, described the bar an idea has to clear there before it gets funded at all:

"sometimes people, for example, will say what it is, but they haven't figured out how to make it a testable hypothesis and you saying, hey, the world might be like this in the future, is like, yeah, you might be right. But that's not what we do." — Astro Teller, Moonshot Factory

The funding structure that follows is the part startups should copy:

"We do not fund things for five years so you can go try it. We're going to fund this for like five weeks, maybe five months." — Astro Teller, Moonshot Factory
"But if you don't have a testable hypothesis, we can't even get onto that flywheel." — Astro Teller, Moonshot Factory

Teller is also honest that even a rigorous process gets timing wrong. On Google Glass: "turns out we were to too early there. But increasingly, it's looking like we were very right." Being right and being early are the same thing as being wrong, on a startup's balance sheet.

The other early-warning mechanism is simply talking to customers before building the architecture. Damien Lewke of Nebulock found his core technical assumptions were badly off:

"I thought we were going to plug into this one system called the SIM and we were going to pull all this data from it. So off base, we ended up needing to plug into an entirely different tech stack." — Damien Lewke, Nebulock
"yes, you have conviction. You want to be obsessed with the problem. You have conviction in your idea, but you do need to check your assumptions." — Damien Lewke, Nebulock

Without that discovery, in his words: "We would have built the wrong solution."

Does going deep in one vertical solve the wrapper problem?

It's the most common escape route, and it works — but it's expensive, and it doesn't transfer. Ben Rudolph described Peregrine's approach as depth-first by design:

"Our whole idea of Peregrine is deeply understand a vertical, so that you can be successful in it and then when we are expanding to emergency operations or expanding to health and human services. We also have to go and do the hard work of learning that vertical, and learning the way we go to market because it's different." — Ben Rudolph, Peregrine

And the warning attached:

"It's very difficult to just kind of copy and paste. Of course, you can take some of those learnings, but you got to do the learning." — Ben Rudolph, Peregrine

That's the honest version of vertical AI. The domain depth that makes you un-wrappable in one vertical is precisely the thing you have to re-earn in the next one. Founders who assume the second vertical is a copy-paste of the first are back to being a wrapper with a new logo — because in the new domain, they don't yet have the context, the liability exposure, or the workflow depth that made them defensible in the first.

Key Takeaways: The Wrapper Test

1. Substitutability is the core question. If a competent team could rebuild your value on the same API in a weekend, the model is the product. 2. Context beats prompts. Cos Nicolaescu's coding-agent analogy is the sharpest test: do you assemble the context that makes the output correct, or hand back one extraction? 3. Don't build on the labs' roadmap. Cos Nicolaescu's phrase "technically interesting, and you can do that" describes most wrappers — and most of them are adjacent to what model companies will ship themselves. 4. Liability changes architecture. Notch built while personally exposed to lawsuits from its own AI's mistakes; that produces verification, not disclaimers. 5. ChatGPT leveled the field both ways. Incumbent spend was wiped out — and so was any advantage you'd have from model access alone. 6. Force a testable hypothesis in weeks, not years. Astro Teller funds five weeks or five months, never five years on a maybe. 7. Check assumptions before architecture. Damien Lewke's plan to integrate with a SIM was completely wrong; discovery caught it before a hard pivot. 8. Vertical depth works but doesn't copy-paste. Ben Rudolph is explicit that each new vertical requires doing the hard learning again.

FAQ: Common Questions About AI Wrapper Startups

Q: Is being a wrapper always bad?

A: Not at day one — plenty of durable companies started as thin layers and built depth underneath. It's bad as a steady state. The risk is that both the foundation labs and your competitors can replicate a thin layer, and Cos Nicolaescu's warning about building adjacent to what the model companies will eventually ship applies directly.

Q: How do I know if I'm just a wrapper?

A: Ask what happens if you hand your customer the raw model instead of your product. If they'd be mildly less comfortable, you're a wrapper. If they'd be unable to reach a correct answer because they lack the assembled context, you aren't.

Q: Does a proprietary dataset automatically make me defensible?

A: Not by itself. What Nicolaescu describes is richer than data — it's knowing how documents relate, what's true about a client, and how to handle nuanced cases. Raw data without that relational context still leaves the model's capability, in his words, "diminished significantly from what it could do."

Q: Should I pick a regulated industry to avoid being a wrapper?

A: It's a strong structural choice, which is why Dan Mishin picked legal explicitly. Regulation creates accuracy requirements, liability, and acute pain — three things thin products can't survive. It also makes everything slower and harder, so pick it for conviction, not just defensibility.

Q: How fast should I test whether the idea is real?

A: Weeks. Astro Teller's standard at X is five weeks to five months to buy one piece of information about whether you're more or less wrong than you thought — not five years to find out.

Sources: Listen to the Full Founder Stories

  • Cos Nicolaescu, Accrual — on context as the real moat, why coding agents work, and the idea maze that ruled out everything adjacent to the foundation labs.
  • Rafael Broshi, Notch — on being personally liable for his AI's mistakes, and why ChatGPT created an even playing field against companies that had spent hundreds of millions.
  • Dan Mishin, Manifest — on explicitly refusing to build a GPT wrapper, and choosing a regulated industry with acute pain instead.
  • Astro Teller, Moonshot Factory — on the testable-hypothesis bar, funding in weeks rather than years, and being right but too early with Google Glass.
  • Damien Lewke, Nebulock — on core architectural assumptions that were completely wrong, and the customer discovery that caught them.
  • Ali Khokhar, Amigo AI — on quitting three months after ChatGPT launched with a five-year category-by-category thesis.
  • Ben Rudolph, Peregrine — on vertical depth as defensibility, and why it can't be copy-pasted into the next vertical.
Listen to the full episodes at pmf.show for the complete stories behind each of these numbers.

Last updated: August 2026

Want more founder stories like this?

Subscribe to The Product Market Fit Show for weekly episodes.

Subscribe Now