Skip to main content

Indie game storeFree gamesFun gamesHorror games
Game developmentAssetsComics
SalesBundles
Jobs
TagsGame Engines

[WIP] Tollens: the quality layer for agentic coding

A topic by qingisbuilding created Mar 03, 2026 Views: 194 Replies: 2
Viewing posts 1 to 3
Submitted (5 edits) (+1)

Tollens: the quality layer for agentic coding [WIP]

One-liner: AI coding lets you ship faster than you can understand what you've built. Tollens keeps you in control.

The problem: You're building something amazing, powered by agentic AI to help write your software as fast as you can dream it. Your codebase is growing faster than your team's ability to reason about what it does, where it breaks, and what happens when it fails. Quality is value to people who matter - your users, your team, your stakeholders. Right now many software teams do not have confidence in what quality they are delivering.

What Tollens does: Real quality work is more than just checklists and test coverage. Quality understanding comes from investigation: exploring what your product actually does, questioning assumptions nobody thought to question, spotting risks that don't show up in a pipeline dashboard. Tollens is an AI toolset that does this adversarial thinking. It flags contradictions, surfaces gaps in understanding, notices when metrics are telling a story nobody's reading, and escalates when a human genuinely needs to make the call - from things as small as "is this UX actually slick to a real user?" to "is this bug a reputation risk we can afford to take?".

Tollens bolts on to your existing AI-native workflow and self-discovers your team's quality context by exploring your codebase and asking questions when it's not sure. Tollens is built around a machine-actionable quality schema: a way of encoding what your team actually cares about (priorities, risks, user expectations) so that AI agents and humans alike can reason about quality based on shared understanding.

Tollens and your QA org: We don't think you can take human judgement out of quality, and we don't think AI is close to changing that. Tollens is your quality team's best assistant, not a standalone quality team. Think of it as what Claude Code is doing for Devs, but for QAs. 

What that looks like depends on who's using it. A junior QA paired with Tollens learns judgment faster, because Tollens asks the hard questions and the junior has to go gather the right input from stakeholders. A senior QA paired with Tollens can leverage much higher impact by amplifying their judgment with Tollens's observation and automation. A QA lead paired with Tollens can confidently generate the holistic quality insights to present to their stakeholders, without worrying about their team not being aligned about the holistic quality assumptions of the product.

Who shouldn't use Tollens: Teams that haven't adopted AI coding tools yet. If you're still figuring out how to actually use AI-generated code in your workflow, your bottleneck isn't quality, it's adoption. Entrenched orgs that are going to get eaten by AI-native newcomers before they can adapt aren't our problem to solve. Tollens is for teams already moving fast with AI and feeling the vertigo.

Why we can have impact: The big labs are raising the waterline rapidly: every base model update can handle more software tasks out of the box. But "software" isn't one problem. There are many genuinely hard tasks within software engineering. For example, five-"9"s (99.999% reliable) telecoms infrastructure, writing embedded firmware you can't patch after deployment, building financial systems where a race condition costs real money, or certifying safety-critical software where someone might get hurt. Out-of-the-box models currently aren't capable of getting the right quality approach to these domains. (I wrote more on this in How fast will AI get better at software?)

We come from one of these domains: three co-founders from the Metaswitch diaspora. We've built trusted software for telecom networks in the cloud. We know what quality thinking looks like when "what if this breaks?" means someone can't dial 999. We've built a quality culture that is just as suited for architecting for five "9"s services on top of three "9"s dependencies as for brutally weighing up exactly what isn't needed for an alpha release to two friendly customers.  

The high-difficulty domains are the summit of the quality mountain. Right now, we're on the foothills. Even on straightforward projects, AI-generated code often ships without considering reliability, supportability, maintainability, security... Tollens scales our judgment with AI tooling, so we can help you get this right whatever your domain.

The timing is also perfect for us to build this capability without being a part of a big lab. The value in AI has shifted to the harness layer: orchestrating models, not training them. (The Harness Layer explains our thinking.)

Next 1-3 months: Super ambitious timeline. Closed alpha launching in ~1 month, targeting ~30 early users: vibecoders with solo noncommercial projects. In ~2 months, design partner trials with real startups, seeing how well we stand up to handling real quality debt in real codebases. By June we want to have validated that Tollens can surface genuine quality insights on codebases it's never seen before, with minimal setup.

Long-term vision: Right now AI-generated code has a reputation as "slop". We want to change that. My hope is that Tollens teaches quality thinking to the software world. My vision is a world where people trust AI-engineered software the way they trust Waymos. Success looks like our way of thinking about quality becoming the default for how teams build with AI.

Resources needed: Mainly feedback on this pitch. "Quality" means something specific to us and something generic to most people, and every time I explain Tollens I end up writing an essay. I want to get to where I can say this in 30 seconds on a call and have the other person get it. Also interested in connecting with CTOs or engineering leads at fast-growing startups who are feeling the quality gap as they adopt AI coding tools - I need design partners to figure out if what we're building works! If you know someone, I'd love to talk. 

Submitted

Updated based on feedback to add sections: - bolts on to your existing workflow - Tollens vs growing your QA team - who shouldn't use Tollens

Submitted(+1)

Great idea! Sounds spiritually similar to https://x.com/VictorTaelin 's "Bend2" concept, you might find them interesting if you weren't already aware of their research. I believe they're hiring, and have an active test question posted that anyone can answer to apply.