Also promised, in theory: Build it and they will come. Agents can run the whole workflow. The product sells itself. Growth is automatic now. It’ll pay for itself.
You’ve come far, fast, and first. But you have your doubts. Trust your doubts. Bring us what you have, even if it’s just an idea, and you’ll have a straight answer within two days. Free. Will it work in production?
Let’s find outPlain words, not a report. You send what you have and where it hurts; Phil reads it himself and writes back with what holds, what doesn’t, and what it’s worth spending next. Sometimes that is nothing, and you will hear that too.
Learn moreSteps 01 and 02, together. Every engagement starts here, because the scope is what opens every door below.
Learn moreAgentic systems built into your operation and wired to your data and tools: research, ops, support, sales. They do real work under human judgment.
Learn moreThe demo already works; production is the part that has to hold. We take your AI-built app the rest of the way and ship it for real.
Learn moreBring us the problem and we build the system: greenfield, with AI and your team, integrated with what you already run.
Learn moreReviews and hardening for AI-built systems, shipped or about to be. If AI built it, someone still has to secure it.
Learn moreSafe, and not because of paperwork. Helping people ship is the practice; the record is twenty-five years long. And frankly, if your idea is worth stealing, someone is already building it. Secrecy won’t win that race. Speed will.
Send it proudly. An AI-built demo is the best first scoping there has ever been: it carries what you know in a form we can read. It’s also probably seventy percent done, and the last stretch is the wall. That’s the part we do.
If we find where it breaks, good. That’s the point. Finding the break is how it gets fixed or made better. And if it holds, you get your go with the evidence to back it. Either way, you stop guessing. AI isn’t the only one with ideas.
Context first, never code or credentials at the free step. The assessment does not rely on AI: the reading and the judgment are ours. We research your field the way anyone would, and what you sent stays out of it. Nothing you send is used to train anything. The NDA arrives with the scope, signed before any code moves.
One straight answer, no drip campaign. But we will check in, the way people do. If you’re moving, we want to hear it. If you’re stuck, we want to help. That’s not a funnel. That’s a connection.
Nothing you send is used to train any model. During the assessment, the only people who see it are us: no subcontractors. We build with agents every day, but the assessment does not lean on them: the reading and the call are ours. Partner companies only ever enter at build, under the same NDA, signed before any code moves. When the work is done, your artifacts are deleted on request.
InTheory is a team of people and partner companies led by Phil Cowan: twenty-five years across product, design, and engineering, building with production fleets of AI agents and human judgment on every call. We build with AI every day, carefully and responsibly, and we have watched every wave of new tech promise everything. The practice is simple: find where AI belongs in your business, and say so when it does not.
25 years at this job · 140+ engagements · the firm since 2011
AI isn’t going anywhere. What matters is where we stand in it. This work was never hands off.
You asked AI if your product was ready. It said yes. You asked again, differently, forty times. Yes, every time. And somewhere in there, a quiet voice said: in theory.
Take that voice seriously. Not because doubt is always right, but because it points at something untested. Doubt isn’t a verdict. It’s a question that hasn’t been run yet.
So we run it. Every scope starts by turning your doubts into tests: name the fear, define what would prove it wrong, go find out. Doubt in, evidence out. Bring it here. It’s on the sign.
The demo is real. It got people leaning in, and that matters. But shipping means surviving what demos never face: a security questionnaire, a traffic spike, a Monday.
Vibe-built apps land about 70% complete, polished enough to look nearly done. The missing 30% is the unglamorous part: auth, data integrity, scale, cost per request, fallbacks, what happens when the model is wrong. Columbia researchers vibe-coded more than fifteen applications and sorted hundreds of failures into nine patterns. The two worst were error handling and business logic.
That’s the part we build. It’s easy to defer from the inside, and it decides everything.
Somewhere in your company there’s a product AI built that nobody can fully explain. It runs. It matters. And if it breaks at 2am, no one’s name is on it.
Ownership isn’t a feeling. It’s a spec: a named accountable person, documentation your team actually runs on, monitoring, an incident plan, and a decision about who answers when the model or the vendor changes underneath you.
Twenty-five years of scoping comes down to one rule: if you can’t support it, you’ve got no business shipping it. Everything we deliver ships with an owner. In writing. No orphans.
AI can make a confident case for almost any answer. Confidence is not evidence. A bluff is exactly that: confidence without evidence. So we don’t argue. We test.
The rubric is written and disclosed: architecture under load, security under attack, the workflow with real users in it, costs at real volume. The same dimensions every time, with tests and thresholds matched to the product’s risks, so the verdict comes from evidence, not from a veteran’s hunch.
And because we also build, we say it plainly: the scope is priced to stand alone, no-go is always an acceptable answer, and the verdict is yours to take anywhere. Judgment you can’t afford to hear isn’t judgment.
Call the bluffWorking isn’t the same as worth it. Plenty of AI systems run perfectly and produce nothing: the demo impressed, the workflow never changed, the money never noticed.
So we measure the boring way: against the non-AI alternative, in real workflows, with real users, counting the errors and what they cost. Before anything gets built, we agree in numbers what paying off means.
If it doesn’t pay, we say so. That’s the cheapest sentence we’ll ever sell you.
You know the loop: one more prompt, one more version, almost right, again. The loop feels like progress. It’s motion.
Done is not a feeling. Ask any artist: putting the brush down never feels right. Done is a line drawn in advance: acceptance criteria agreed before the iteration starts, and kill criteria for when to stop entirely.
We draw that line before we begin, and we hold the work to it. Ours included. Done is a gate, not a wall: the product ships, gets owned, and earns its next round. Closure ends the lap, not the road. That’s the job.
The tools have changed. Someone still has to decide what will work in production, and be the first to say no.
Twenty-five years, 140+ engagements. Creative direction on Disney, Pixar, and Universal film campaigns. ShapeShift and SALT Lending launched out of our own office, fractional CTO through SALT’s launch. Fund rails for WallStreetBets. Today the same work runs on a team of people and partner companies, plus a fleet of specialized agents on an agentic OS.
Denver, since 2011. Through the whole crypto cycle and into this one. The frontier days of crypto and the upheaval around AI now are the same thing at a different speed, and we have worked through both. Plenty of shops chased each wave and are gone. We are still here.
No charge and no obligation. Code only ever moves later, under NDA, with the scope.
Let’s find outYour procurement team will recognize the shape and your engineers will recognize the substance. Here is how it gets made and what lands on your desk at the end of it.
Sessions with whoever actually holds the problem: you, your engineers, and the person who has to sign for it. We want the wall you hit, the thing you keep deferring, and the part nobody wants to own. Not a requirements list. Requirements come out the other end, not this one.
If your prototype exists, we go through it with the people who built it, and we ask what they’d change if nobody was watching. That conversation is usually where the scope really starts.
Against a written rubric you get to read first: architecture under load, security under attack, the workflow with real users in it, costs at real volume. Where we need access to your systems we ask for it up front, in writing, and never at the free step.
Two jobs. It validates the direction before anyone spends a build budget on it, and then it guides the production, milestone by milestone, for whoever does the work.
It stands alone and it’s yours to take anywhere, including to someone else’s build team. We’d rather you built it with us. But a scope that only works if you hire us isn’t a scope, it’s a pitch.
No-go is always an acceptable answer.
The scope runs under NDA, signed before any code moves. Nothing you send trains any model, and your artifacts are deleted on request when the work is done.
Let’s talkHaven’t had the free assessment yet? Start there, it costs nothing. Rather read this in full? The anatomy, on its own page.