18 Years In, and We Don’t Hire Developers Anymore — Vadim Peskov, Diffco

AI agent development is what Diffco does now, but the company is 18 years old — it started as web development in 2008, moved through mobile, and arrived here. In Episode 7 of the NoRobots Podcast, Vlad asks founder Vadim Peskov what clients actually pay for when agents write all the code, why project costs fell while rates went up, and who is liable when AI-generated code leaks data.

Eighteen years of changing the answer

What a client needed in 2008 has nothing to do with what they need now, and Vadim credits the company’s longevity to changing processes repeatedly rather than committing to one mantra. Asked whether there was a master plan, his answer is dry: they were stupid enough to start in 2008 and continued by mistake.

The AI work goes back further than most people assume — computer vision and ML projects from 2016. Selling it then was mostly education, because nobody understood what the technology could do, and some clients expected 100% recognition accuracy. Back then 92 to 95% was a success; now 99-plus is normal, better than the humans it replaces. The bigger shift is that they used to train models and invent things that didn’t exist. Today there are thousands of models to test in parallel, and it’s a question of budget rather than invention.

One hundred percent

Asked what share of delivered code is written by AI, he doesn’t hedge: for projects delivered in the last six months, 100%.

The exceptions are specific. Compliance-heavy systems where the stakes are high stay closer to manual with AI assist. Pixel-perfect front-end work isn’t there yet — if the design has to match exactly, full autonomy is out. He expects 90%-plus even on compliance projects by year end, with human oversight layered on top of the AI oversight. His line on where that matters: you don’t want to ship a banking system moving millions of dollars a minute with AI.

So what is the invoice for?

Nobody ever paid Diffco per line of code, and he points out that “you wrote another 500 lines, here’s money” was never how the business worked.

What clients pay for is the system that turns an idea into a working product: the specification, the agent engineering, the memory management, and knowing what the technology can’t do. Ask an AI to deploy on AWS in the best fashion and it will do something — if you don’t understand what it did, that something is useless.

The part he emphasises most is partnership: being the firm that occasionally says this is a bad idea and here’s a better way. He points at a long graveyard of products built over thousands of hours by founders who never tested the concept with a paying customer, and only discovered afterwards that what they’d built wasn’t what anyone wanted.

Rates up, client costs down three times

Diffco raised its rates substantially. Client costs still fell roughly threefold over three years: a project that would have cost a million dollars now lands under three hundred thousand.

The gap between what’s possible and what’s permitted shows up inside a single enterprise client. One department does pixel-perfect work, manual with AI assist. Another department of the same company adapts prototypes and ships fast. The delivery rate in the second is 20 to 30 times the first. Same company, desks next to each other, different appetite for risk.

Clients arrive with prototypes now

Roughly 70 to 80% of incoming projects come with a prototype, often built in Lovable. Vadim welcomes it, especially when the prototype actually launched and gathered real usage data, because what surfaces then isn’t bugs but user-flow problems — and those experiments cost the client almost nothing.

Their approach is guardrails rather than gatekeeping. Clients can change copy, images, and flows themselves. Touch the database structure and a human or a more capable agent reviews it first. By his estimate, 90% of what founders and marketing teams want to try can be handled that way.

Which moved the bottleneck somewhere unexpected. Agents now produce work faster than humans can approve it, so inside enterprises the approval flow — not engineering capacity — is the constraint.

What vibe coding doesn’t solve

His objection isn’t to vibe coding but to the belief that half a page of specification yields a complex product. A brief that short means many conversations before anything gets designed; on some projects they’ve had literally hundreds of calls with a client to pin down requirements, because the decisions are complex and the regulation is real.

Build-versus-buy has genuinely changed, though. A tool Diffco uses announced a move to an enterprise plan that would take their bill from a few hundred dollars to roughly fifty thousand a year. They found an open-source alternative within minutes and are migrating within days, and may fork it into their own harness and contribute back. Five years ago he would have asked why anyone would bother.

But he doesn’t extend that to everything. For another tool — effectively several products stitched together — rebuilding is possible and pointless, because maintenance would cost more than the subscription. He mentions a company that rebuilt about 5% of Salesforce in two months with three developers and saved a million a year in licences; it works for them, and it’s still a distraction for most. His framing: software is never finished, and if you stop at version one you die. Price the maintenance, not the build.

Who is liable when it leaks

On AI-generated code and security, the legal answer is that the client ships the product and carries that risk, while Diffco insists on at least some security testing — a request clients rarely refuse.

Then the reality check. No system is ever fully secure; if you want one, don’t launch and put the computer in a safe. Across 18 years, he attributes roughly 95% of incidents to employees doing something they shouldn’t have and the remaining 5% to systems clients declined to update. And the multiplier people don’t price in: if AI writes code a thousand times faster, the surface for mistakes grows accordingly.

His practical floor is cheaper than most expect. A standard pen test runs a few thousand dollars, broader security testing perhaps five, and doing either puts you ahead of most teams shipping today. With no budget at all, pointing a capable model with the right setup at your own code and network will still catch the majority of it. The rest — disaster recovery, backup strategy, incident response — is a company question, not a developer question.

The junior problem, unanswered

Junior hiring is down 30 to 50%, and asked how anyone becomes senior in five years if nobody hires juniors, Vadim says honestly that he doesn’t have a good answer.

What he can describe is what replaced the role. Diffco doesn’t hire developers; it hires engineers who design systems that build themselves. The interview filter is abrupt — introduce yourself as a front-end developer, or a back-end developer, and they’re not hiring. Their technical product managers build more complex systems than many developers do, because they write better specifications. He’s equally blunt about seniority: fifteen years of experience counts for nothing if you aren’t working this way now.

His advice for someone starting out is to experiment relentlessly and never deploy blindly. If you ask an AI for an architecture and it returns Docker swarms across multiple clouds for an application with a hundred customers, something is wrong — a single server is fine. Knowing enough to challenge the machine is the job.

Growing in a different direction

Diffco has fewer developers than it used to and is growing faster than ever. Vadim expects the developer headcount to keep shrinking in favour of architects, while the work that remains is understanding what a client actually needs and walking them through the requirements.

Meanwhile the agent count goes the other way: 30 to 40 running across their systems today, heading toward hundreds or thousands. Humans are still needed to decide where to point them — without that, he says, it’s a journey into nowhere.

Find Vadim on LinkedIn or at diffco.us, and watch the full episode on YouTube. Diffco also posts open roles on our job board. For another take on where AI stops and humans pick up, see our episode with James Zammit of Roark.

Three Products in Two Years — Harsha Gaddipati, Slashy

An AI email client is Harsha Gaddipati’s third product in two years. The first two both hit ten thousand users in their opening week, and both failed anyway. In Episode 6 of the NoRobots Podcast, Vlad asks him what went wrong twice, why email turned out to be the right thing to own, and what people actually automate once they have the tools.

Ten thousand users, twice, and still no business

The first product was a website builder, launched right as Lovable was taking off. Social posts went viral, ten million views across platforms, ten thousand users in week one. You would assume that is a company.

It wasn’t, because the use cases were one-offs: a resume, a business site, something for a partner’s anniversary. People finished the task and never converted off the free trial. The lesson was that you need to own a workflow that repeats.

So they built a general AI agent. Same result on signups, same problem underneath, but with a sharper edge. Power users loved it and described it as the thing they reach for when ChatGPT or Claude fails. Harsha is blunt about what that means: if your market is what the big models can’t do, your addressable market shrinks every week, and you eventually die not because the product is bad but because that entry point is already won.

Why email

His framing for the third attempt: there are only about six places the work you want to hand off actually arrives. Email, messages, Slack, your to-do list, meetings you take in person, and your own anxieties.

Of those, email is the easiest to disrupt, because it lacks the network effects the others have. Slack needs the whole organisation to move. iMessage needs everyone you talk to. Email can be adopted one person at a time.

The trade-off is honest: getting someone onto a new email client is harder than getting them to try a chat tool. But losing them afterwards is also much harder. His summary of the priority — when you’re trying to build something enormous, retention matters more than acquisition.

The forty-hour sessions

Reports of forty-hour building stretches are accurate, but the framing is not what people assume. It’s never a decision to work forty hours. It’s a milestone: first users tomorrow, so what has to be true by then.

The constraints — sending works reliably, labelling works — are usually met in the first six to eight hours, especially with AI doing the typing. The remaining time goes into walking the user flows and fixing the parts that don’t feel right, which is a different bar than passing a test. The two founders swap: one writes a feature, the other tests it.

His reasoning for why that matters: email is a workflow, and in consumer products the user has to feel some joy getting through it, not merely complete the task.

Selling against a switching cost

Nobody pays for an email client if their life isn’t busy, so the only prospects worth talking to are the people with the least time to switch.

The pitch that works isn’t “your current process is broken” — a successful person reasonably believes their process is fine. It’s an outcome question: do you feel you have enough time in the day, and are you hitting your goals? If not, and nothing else is going to change, why not spend thirty minutes on setup for the chance of saving an hour a day.

Beyond that, it’s word of mouth, which he considers the number one way anyone tries anything, plus workshops that put the product in front of people who already trust someone using it.

How the drafting learns

Most tools treat drafting as an isolated action. Slashy treats it as one action in the ongoing life of an agent that starts learning the moment you sign up, which is why the output drifts closer to your voice over time.

What feeds a draft: your past emails, a lookup of who the recipient is and what their site says, your previous emails with a similar goal — a contract negotiation pulls your other contract negotiations — plus your calendar for anything involving times. And crucially, the diffs between what it drafted before and what you actually sent. Models carry their own biases; without seeing the correction, they never converge on you. On top of that sits a memory layer indexing facts about the user, down to food preferences.

Flat pricing, on purpose

Pricing moved from credits to a flat thirty dollars a month, and the reasoning is a bet about where inference costs are heading.

If you charge per AI call and the cost of intelligence trends toward zero, you have built a business that shrinks as the technology improves. He points at coding tools as the illustration: a user switching from an expensive model to a cheap one takes most of that revenue with them. For a company aiming at real scale, he considers pure usage-based pricing a poor long-term structure.

The features nobody expected, and one nobody used

Two surprises came from customers rather than the roadmap. The Slack and iMessage bots shipped with no expectation that anyone would care — and became a favourite among founders and growth leads, whose work mostly happens outside the inbox. Instead of keeping Gmail open and refreshing it, they get texted when something important lands, which removes the context switching more than it saves clicks.

The second: people started using it as an executive assistant for scheduling, which it was never designed to be. The team knew it could draft, prep calls, label, and track follow-ups; customers worked out the rest.

The failure is more interesting. They shipped tab-complete — autocomplete for the next sentence — and nobody used it. People keep asking for the feature to this day, having ignored it when it existed.

What to automate first

Asked what a ten-person company should automate on day one, his answer is deliberately unglamorous: not replies, but reminders that you forgot to reply.

His view is that people badly underestimate how many important emails they drop in a week — a missed follow-up, a customer left waiting — often without receiving much mail at all. He sets a hard bar for non-technical roles: unless you’re in a meeting, there’s almost no excuse for taking longer than ten minutes to reply.

For his own mail, essentially everything starts as an AI draft and gets tweaked, usually for things software can’t do, like attaching a screenshot.

Cold email that actually lands

His advice runs against the usual instinct. Personalising what your product does for someone rarely moves the needle — if they have the pain, they were already interested; if they don’t, listing features won’t create it.

What moves the needle is evidence you looked into the person. Find a genuine interest, connect it to something you share, and lead into the product from there. It reads as written for them, it’s hard to get wrong, and the data is easy to find, whereas guessing what a company needs is easy to get wrong.

What’s next

The stated one-year goal is narrower than the ambition: brand recognition. Today, “email for high performers” makes people think of Superhuman, Shortwave, or Fyxer. He wants that reflex to be Slashy. His view is that once the brand sits at that level, most other problems solve themselves.

Find Harsha at slashy.com, and watch the full episode on YouTube. For another founder’s take on building a product people keep paying for, see our episode with James Zammit of Roark.

Trying New Things Without Breaking What Works: Todd Anthony of Pinwheel Agency

Todd Anthony is Partner and Executive Creative Director at Pinwheel Agency, which he has been running for over 12 years – a rare run in the agency world. Before Pinwheel there was another agency, a partner that didn’t work out, and a lot of lessons learned the hard way. We asked him how he tries new tools and ideas, what he does when they don’t work, and what he still wants to try.

What’s something new you’ve tried recently that changed how you work?

It really feels like “how we work” changes constantly these days. I was in a meeting with our client Stripe recently to discuss creating a graphic for one of the breakouts at their annual conference, Stripe Sessions. I had been typing careful notes as the client described what she was looking for. And I had a kind of aha moment. I thought, what if I just feed these notes into Claude right now and have it generate a prototype or rough sketch of how we understand the idea, so that we could ideate on it in real time. And I think that basically changed, for me, the nature of these complex graphical projects. Clients don’t have a lot of time these days for iteration, and a lot of the collaboration is happening async. Being able to do this IN the briefing meeting WITH the client not only speeds the process, but it improves the quality of the collaboration.

How do you test new ideas without risking client outcomes?

Well, for starters, we don’t charge for them. In fact, we’re testing a new idea right now. Working with a technology partner, we’ve built a brand health diagnostic tool that collects signals from all over the internet to basically hold up a mirror to a brand and show them how everyone else sees them. It points out things they’re doing and not doing that tend to cut against the brand image they’re trying to project out into the world. And it points out inconsistencies and the ways in which they can/should bring their brand into alignment with itself, better reach their buyers, and drive higher revenues.

So we tried it out on one of our client brands, but didn’t send it to them. We just worked on it through what we perceived to be their perspective. We knew their marketing programs, their internal challenges, and the way that they were structured. Therefore, we knew there were things in the report that they wouldn’t care about, and some things we knew they couldn’t do anything about. There were also some areas in the report that, upon further reflection, weren’t so clear. After fine-tuning the tool, we finally offered to present their brand health report to them – free of charge. They seemed enthusiastic, and the meeting went well, but they left the meeting not quite sure what to do with the information. So we’re now making the tool more actionable and useful to marketers. It’s been a fun process that only cost the client an hour of their time. And while it’s certainly an investment on our part, we think it’ll be highly valuable to current and future clients.

What’s an experiment that failed but was still worth it?

Oh my gosh, so many beautiful, glorious failures. I had an agency before I owned Pinwheel called Minty Fresh, and I took on a partner early on. At first, I thought we both wanted the same things, but then it became clear we had completely different visions for what we wanted the agency to become. And his style was absolutely NOT compatible with mine. He was more showy and braggy, and tended to stretch the truth quite a bit. Not to pat myself on the back, but I’m the exact opposite. We were 50/50 partners, and things eventually came to a head. I knew we were going to have an ugly fight over control, so I offered to sell my half to him – just as a way to avoid some of that pain and anguish. We still had to have lawyers intermediate the whole thing, and it did get pretty ugly at the end, but I gleaned some extremely valuable lessons from the whole experience: the importance of vetting your partners very thoroughly, that you should always maintain a dominant share, that you don’t need a lot of the overhead you think you need, and so on.

And, if I’m being honest with myself, I have to acknowledge that I learned a lot from him as well. The notion of having bigger ambitions, of not being afraid to self-promote (though absolutely DO NOT stretch the truth), of realizing that I’m actually really good at what I do, and so on. And all of that learning has helped me navigate the Pinwheel journey for over 12 years now – which I’ve learned is fairly rare in the agency industry.

How do you decide when a new tool, trend, or technology is worth integrating into your workflow?

We don’t have a rubric, but when a tool seems to have enough traction to gain the attention of a significant portion of our agency and client brethren, starts showing up in conversations and newsletters for a period of time, and sounds like something that will help us, we give it the ol’ college try. Some stick, some don’t. Granola is an AI notetaking app that we love very much. When AI came onto the scene, I started using a product called PrettyPrompt that automatically improved prompts – making the outputs far more useful. The models have improved since then and I don’t need it as much, but it still comes in handy. We used Trello at one point. Didn’t love it. So we switched to Airtable after we heard from other agencies that they loved that product. We also have a section of our twice-monthly marketing newsletter, the Spin, called “Tool Joy” where we review new tools that we think our marketing partners might benefit from. And I follow Product Hunt so I can stay on top of the incredibly swift pace of digital innovation that’s changing the way we work on an almost monthly basis.

As a side note, we like to think that we’re highly adaptable, and we are. But there is a cognitive load that’s getting heavier and heavier to carry due to the increasing pace of change. At some point, the cost of changing will be higher than the benefit that would be gained from that change – even if there is a clear value proposition. And that’s because we’re basically slightly smarter apes.

What’s something you want to try but haven’t had the chance yet?

I was scheduled to go skydiving about 30 years ago in the former Czechoslovakia with a friend of mine. It was going to be in a Russian plane with American parachutes, which is a configuration far superior to its opposite. We did the training. I started getting butterflies. We were about to go up, but then the wind picked up and they had to postpone for a day.

Ready to go the next day, we had a miscommunication and my friend didn’t show up to the pick-up spot. I tried finding buses to the airport, but couldn’t quite sort that out (this is pre-internet, pre-smartphone). So I never got to do it. And as I’ve gotten older, I’ve lost my nerve a little. Isn’t that funny how that works? You’ve gotten far more out of your life (in my case I’ve been married, had two kids, owned two companies, etc.), and yet you’re less inclined to risk the remaining time. Should really be the other way around, don’t you think? In any case, I’d still jump if there was an enthusiastic friend instigating it.

Final Thoughts

Todd’s approach to new things is simple: try them on your own dime, keep what works, and be honest about what didn’t – including a partnership that ended with lawyers. After 12 years at Pinwheel, that mix of curiosity and caution seems to be the thing that keeps the agency going.

Explore More Agencies

If you’re an agency, designer, or startup looking to boost your visibility, you can join Norobots and become part of our curated network of trusted businesses. If you’re a brand or client searching for the right partner, our platform helps you discover agencies, designers, and startups you can rely on. Browse the listings and find the right fit for your next project!

How a 6-Person Startup Landed Google & AT&T — James Zammit, Roark

Voice AI agents are moving from novelty to infrastructure, and the companies building them have a reliability problem nobody talks about. James Zammit co-founded Roark (YC W25) to solve it — and did it with a team of six that counts Google and AT&T among its customers and processes millions of call minutes a month. In Episode 3 of the NoRobots Podcast, Vlad asks him how a company that small lands enterprise clients, and where he thinks voice AI still has no business picking up the phone.

Two startups that were too early

Before Roark there were two others: a music collaboration app in 2015 and a chatbot company in 2018, back when chatbots were universally disliked and the underlying technology simply wasn’t there. James is clear-eyed about why they didn’t work — wrong timing, no domain expertise, and founders who were early in their careers.

Two corrections matter more than the rest. The first is speed. One earlier product took a year and a half to reach its first version, because they wanted it perfect. By launch, either someone had beaten them to it or the thing they’d polished turned out to be something nobody wanted. He borrows an analogy from Michael Seibel: to find a leak in a pipe, you can inspect every inch by hand, or turn on the water and let the leak announce itself.

The second is that they used to believe a great product sells itself. It doesn’t. Distribution and sales matter just as much, and that lesson took a decade.

What Roark actually does

Anyone shipping a voice agent hits the same wall: you change a prompt, and you have no idea what you broke. The only way to find out is to call your own agent, over and over, through dozens of scenarios. James and his co-founder hit this themselves building an agent for a dental clinic, where patients kept getting stuck in loops and failing to confirm insurance.

Roark started as analytics — “the Mixpanel for voice” — and grew into two things: simulation testing that runs an agent through hundreds of scenarios, personas, and accents in parallel inside your CI/CD pipeline, and post-call analysis that scores every real conversation on dozens of metrics and points at what to fix.

How six people landed Google and AT&T

No ads. The mix was outbound to teams building in voice, referrals, published content, and — for some of the larger names — inbound. Then the ordinary sequence: demo, pilot, contract.

The unglamorous truth underneath is that none of that survives a bad product. James keeps returning to the YC line about building something people want: simple to say, hard to actually do.

Twenty agents instead of twenty hires

Roark runs more than twenty internal agents. Three examples give the flavour.

The changelog agent reads the week’s pull requests, drafts the changelog, and takes its own screenshots. Monday morning someone reviews it and sends. Four or five hours of a PM’s week become thirty seconds of attention.

The on-call agent picks up CloudWatch alarms, investigates, and either opens a pull request for a human to review or writes up what it found. The broadcast agent solves a small, real annoyance: Roark keeps a shared Slack channel with each larger customer, and announcements used to mean pasting the same message a dozen times. Now it goes into one channel called Broadcast and fans out, picking up new customer channels by naming convention. Three prompts replaced a job that other companies pay a tool to do.

None of these are impressive alone. That’s the point — they’re the boring work that quietly justifies the next hire.

The unglamorous half: infrastructure

The other half of staying lean is refusing to let engineers lose days to setup. Everything is infrastructure as code, using Pulumi and SST. Every engineer gets a production-like environment with one command, with preview deploys, end-to-end tests, and alerts on everything.

James has worked somewhere that editing a single page meant three days of fighting to get data running. His argument: with today’s coding tools, setting things up properly takes no longer than doing it badly.

Where voice AI should stop

Appointment scheduling, reminders, prescription refill nudges — all fully automatable today, in his view. His line for what shouldn’t be automated is risk-based: if getting it wrong costs a life or a six-figure deal, a human still picks up. Anyone calling in genuine crisis belongs in that category.

He’s equally candid about what still breaks. A bad phone line plus background noise — a human decodes that better. People switching languages mid-sentence. Interruptions, where the hard part isn’t handling them but stopping the agent from rudely cutting the customer off. And pronunciation: ask an agent to read back an email address and watch it mangle the spelling. Entire companies exist to fix just that.

The architecture is shifting underneath all this. For the past year most production systems were cascade models — separate speech-to-text, LLM, and text-to-speech components stitched together. Over the last six months more of Roark’s customers have moved to speech-to-speech models, which handle interruptions so naturally that callers often don’t notice they’re talking to software.

Should the agent admit it’s AI?

James expects regulation to answer part of this, and Roark follows whatever the local law says. His own position is a hybrid: booking a table, nobody cares who answered. Discussing a prescription, probably say it. Finding exactly where that line sits, he admits, is genuinely hard.

His broader thesis is blunter: anyone with a publicly facing phone number will eventually be a voice agent. And there’s an argument that the machine is sometimes better — a legal office that closes at five loses the customer who calls at six, and an agent doesn’t.

Malta, San Francisco, and what YC actually reads

James and his co-founder are both Maltese and moved to San Francisco. His advice on geography: move to where your customers are — New York for finance, London for parts of Europe — but don’t let location stop you from starting. If you won’t relocate, commit to being there a couple of weeks a quarter.

On YC applications, his read is that where you’re from matters far less than where you’ll be during and after the batch. And what they’re really testing isn’t the idea but how you got to it: did you talk to twenty people, do they all have this pain, what are they doing about it now, will they pay. Building blindly on a hunch is the thing that fails.

For the rest of the conversation on where voice AI is heading, see also our episode with Elad Hefetz on how AI is reshaping search and discovery.

Find James at roark.ai or on LinkedIn, and watch the full episode on YouTube.