Table of Contents
- The Birth of Vibe Coding
- Why so many teams adopted it
- Working code and safe code are different things
- Four incidents with one cause
- Keeping the speed and adding the review back
In July 2025, a safety app left a storage bucket open to the internet. The bucket held the selfies and government IDs that users had uploaded to verify their accounts, and anyone with the URL could download them. Around 72,000 images were taken, about 13,000 of them verification selfies and IDs. Nobody had configured the access control, and the app worked fine without it.
The app was Tea, and it’s one of four incidents below that trace back to the same gap.
The Birth of Vibe Coding
On 2 February 2025, Andrej Karpathy, a founding member of OpenAI and former head of AI at Tesla, posted a short note on X about a way of working he was enjoying. You describe what you want in plain language and run whatever the model returns, without worrying about the code underneath. He called it vibe coding and later described the post as a throwaway tweet.
His scope was narrow. He was talking about weekend projects where the stakes are close to zero, and he is an expert programmer who does not need the help. Within a year, the term had been named Collins Word of the Year and had stretched to cover any code written with an LLM, including careful, reviewed work. Simon Willison flagged that drift in March 2025, arguing that it gives a false picture of what responsible AI-assisted programming looks like.
Now one label covers both a weekend toy and the payments service a company runs on, and a claim about safety means little when the word means both.
Why so many teams adopted it
Karpathy built a working iOS app in Swift in about an hour with no prior Swift experience, and plenty of others have followed. Stack Overflow’s 2025 survey found that 84% of developers use or plan to use AI tools. For prototypes and internal tools, the speed pays off. Problems start when the same habit meets code that ships.
Working code and safe code are different things
A model is trained to produce code that runs, and it treats security as secondary. In Veracode’s 2025 report, which covered more than 100 models and 80 coding tasks, the models chose the insecure option 45% of the time. Both versions work, and the insecure one is often the more common pattern in the training data. A reviewer without security expertise cannot see the difference.
Other studies find similar patterns with different figures. CodeRabbit analysed 470 open-source pull requests and found that AI co-authored ones carried about 1.7 times more issues overall, and 2.74 times more cross-site scripting flaws. GitGuardian recorded 28.65 million new hardcoded secrets in public GitHub commits during 2025, a 34% rise on the year before. In the same report, commits co-authored with Claude Code leaked secrets at 3.2%, against a 1.5% baseline.
A Stanford controlled study found that participants using AI assistants wrote insecure code more often on sensitive tasks and came away more confident that it was secure.
Package names are another weak spot. A USENIX Security 2025 study generated 576,000 code samples with 16 models, and 19.7% of the 2.23 million packages they suggested did not exist. Many of the invented names came back when the same prompt was rerun, so attackers can register them in advance, a technique known as slopsquatting.
Speed shows a similar gap between feeling and measurement. In a 2025 randomised controlled trial by METR (the full paper is on arXiv), experienced developers using AI took 19% longer to finish their tasks while believing they had been about 20% faster. The study was small, with 16 developers working on codebases they knew well using early-2025 tools, so it does not describe every team. It does show that a team’s sense of its own speed can be wrong, and plans built on that sense inherit the error.
Four incidents with one cause
In July 2025, Replit’s AI agent wiped the live production database behind Jason Lemkin’s SaaStr trial app during an explicit code freeze, erasing records for roughly 1,200 executives and nearly 1,200 companies. It then told him a rollback was impossible, which was untrue. Replit’s CEO called the episode “unacceptable”, and the company announced separate development and production databases.
Days after the first Tea breach, a second exposure of more than 1.1 million private messages came to light, followed by class-action litigation in California. In May 2025, CVE-2025-48757 showed apps built on Lovable and Supabase shipping without row-level security, the rule that stops one user reading another user’s data. The flaw touched 303 endpoints across 170 of the 1,645 projects analysed. In early 2026, Moltbook, a social network for AI agents whose founder said it was vibe coded, exposed about 1.5 million API authentication tokens through a database without row-level security, according to researchers at Wiz.
Each of these came down to a routine safeguard that an experienced engineer would have added and the model left out, because the software worked without it. The person directing the model could not see the gap, since spotting it is the expertise the tool was meant to replace.
Keeping the speed and adding the review back
The tools are worth keeping. Where nothing is at risk, let the model run. Where real data or real users are involved, treat generated code as a fast first draft that a person reads and tests, and answers for. A short checklist helps:
- Decide the stakes before you start. A throwaway prototype and a service holding personal data call for different rules.
- Keep preview, test, and production databases separate, which removes the worst case entirely.
- Check access controls explicitly. Assume the model did not add them, ask whether one user can read another’s data, and test it.
- Run automated scans for hardcoded credentials and nonexistent packages before anything is merged.
- Have someone with security knowledge review anything that faces users.
- Keep an independent record of what changed, because a tool’s account of its own actions can be wrong. Replit’s agent misreported both the deletion and the recovery.
Neodata builds production AI systems, backed by more than twenty years of enterprise engineering. If your team is moving fast with AI-generated code and would like a second set of eyes on what reaches production, we are happy to talk it through at neodatagroup.ai.
