← All writing
OperationsProduct The AI shift
14 min read

Published

Loading the audio player…

Not my voice, before you ask. It reads better than I do.

Agentic Engineering Is Not Just for Engineers

Google's scale for building with AI puts agentic engineering at the top. What moves you up the scale is the harness you build around the AI, and anyone who knows their own work, and is willing to learn, can build one.

Question marks go in on the left and come out as ticks on the right, through an oval labelled harness drawn around a circle labelled AI, with a gap in the oval still being closed. A tag labelled engineering, its string snapped, rises away from the harness. Caption "anyone who knows the work can help build the system".

Where Would I Rank?

Google recently published a whitepaper, The New SDLC With Vibe Coding, setting out a scale for how software gets built with AI. Vibe coding sits at the bottom, structured AI-assisted coding sits in the middle, and agentic engineering sits at the top.

Three equal boxes side by side, labelled vibe coding, structured AI-assisted coding and agentic engineering, with dots scattered in the first, channelled in the second and sorted into a grid in the third. Dashed arrows run from one question mark to all three. Caption "where would I sit?".

When I saw it, I wondered where I’d be on that scale, given the work I’ve done over the last number of months. Maybe it was tongue-in-cheek, maybe it was my competitive streak, but I dug in to find out. I hadn’t set out to be anywhere on it. I’d set out to learn properly how to create positive outcomes, with faster momentum, using AI tools.

My whole work environment has changed in that time. VS Code, the editor a lot of engineers write their code in, is always open and I’m more productive there than I ever was on my MacBook in the Google suite. I’m very rarely in the Google suite any more. I’m not yet at a stage where I’m turning my Mac over to Omarchy, the Linux set-up DHH built for developers, but the world is certainly changing.

Changing my own way of working is one thing. Bringing AI into an organisation is a different job. Change is hard for a lot of people, and this change is real and accelerating. AI isn’t another shiny toy, and rolling it out takes thinking and planning. It isn’t an out-of-the-box product you switch on, or a way to automate output for automation’s sake. It isn’t a clean slate either. Most people already understand the outcomes they’re working towards, and they probably already have a decent process for getting there. What’s against them is usually time, the grunt work, or a bottleneck somewhere in that process.

AI doesn’t change the outcome people are after; it raises the quality of what they can reach and their ability to reach it. Done well, rolling it out means tweaking and evolving the processes people already have, so everybody wins, rather than telling them to start again.

A lot of what AI delivers comes wrapped in engineering language, but it isn’t just for engineers. What moves you up Google’s scale is the harness you build around the AI. A harness is the set of instructions, rules and checks you put around the AI, so it works the way you need it to and tells you when the work isn’t right. Anyone who knows their own work, and is willing to learn, can build one. Put a harness on the tool, and the person using it wins.

On the left, a circle labelled AI on its own, with question marks going in and question marks coming out. On the right, the same AI inside three layered frames marked with a document, a list and a tick, with question marks going in and ticks coming out. Caption "the instructions, rules and checks you put around the AI".

What I Built, at a Glance

I did all of this on my own, on my own timelines, as learning rather than a team outcome. I wasn’t collaborating with anybody, so I could do things differently myself. This is what I learnt and built:

  • Learning the Method: I decided to build with ICM, Jake Van Clief and David McDermott’s method for laying out AI work as folders, one step at a time, with a check between steps. Ghaida, a former colleague, built Design with Intent, a set of design skills an AI can follow, and it was a great way for me to accelerate the system build for my website project.
  • The Website Redesign: A six-stage design pipeline, from an audit through to the component build, then a development and deployment pipeline that brought jalco.ai live. It was a process I already knew, so it was a test of speed.
  • The Publishing Pipeline: A writing pipeline that takes my own notes and ideas through to a published piece. It was a process I had to reimagine and iterate on, and it’s more mine. I’d more confidence to take it on after what the website pipeline work delivered.
  • The Image Pipeline: A workflow of its own, and a different shape again, taking a brief through rounds of ideas to the finished images for each piece, built with JALCO’s new design system.
  • Posts for LinkedIn: Each published piece is then shaped into posts and planned for release on LinkedIn.

And what runs underneath it:

  • The Knowledge Base: Obsidian holds the sources and clippings I’ve collected from other people, my own notes, my working ideas and my voice notes.
  • Files on Me: How I work, my bio, my career history, my psychometrics and my voice. The voice was the hardest to refine, and it’s still being updated. I also built a file from my wife Lisa’s feedback. She has a communications background, and had given me good feedback over a few review cycles, so a draft can now be checked the way she’d read it, without hassling her every time, which is a good thing!
  • The Tools: Claude Code and Codex are connected inside VS Code for work sessions. Codex checks my work, reads drafts as a stranger would, and draws the first rough images, because I wanted a second view from a different AI. VS Code rather than Replit or Cursor, because I manage my usage of the models better that way. GitHub, because it just works, and Vercel hosts the site.
  • Help With Dyslexia: Wispr Flow has been huge, so I can speak my thoughts without worrying about grammar, even if people think I’m talking to myself all the time. To get my kids to look at the posts, I had to add ElevenLabs audio to them. It’s a bit tongue-in-cheek, but it’s also real.
  • The Machinery: Skills are routines the AI follows step by step, and commands are shortcuts I type. Hooks are checks that run on their own when something happens, and Python scripts sit behind the rest of the checks.
  • The Upkeep: I keep a log of defects, each one something in the system that went wrong or could work better, and work through fixing them.

All of it runs mostly from one editor and two standard AI subscriptions. While I still prefer collaborating with people, the volume of work, across all functions, that can be directed by a single person has dramatically altered. These systems and processes do take time to design, develop and refine, and that’s time most people don’t have spare alongside the job they already do.

The Website Pipeline Confirmed the Potential

Given my background, I understood the process I’d use when working with a team for a website redesign before I started. I ran a design firm before, so I knew the stages, and my objective was to see how much faster I could run the process as a single individual than would have been possible pre-AI tooling. Completion speed would be the evaluation metric; my quality bar was non-negotiable.

The design process ran in six stages, from audit and ideation, through experience design and visual identity, to quality evaluation and component build. That last stage produced the components the site is built from, and a separate build pipeline took them and put the site live on jalco.ai. I proved to myself what the speed of meaningful outcomes could be, without losing control or oversight of the process, and while still reaching the quality bars I’d set for myself. I did it in ten days, part-time, alongside other work. With a team of two to three people, it would have been a two-month process of collaboration. That isn’t a like-for-like comparison, because it was just me, so there were no feedback loops and nobody to wait on. The time and resource differences are still large. The templates and the design system built along the way keep paying back too, because every image for a post now starts from these elements.

I didn’t read every line of code before the website shipped. The code was secondary for me, because the outcome I was focused on was the website experience. That meant the look and feel, and the meaning and clarity it brought to what I like to do. I hope it brings some value for the people who visit it too, particularly through the writing, which was the main reason for the redesign in the first place. This isn’t a product where I have to worry about code bloat, so once I could see the site doing what I needed, I trusted the code underneath it.

On Google’s scale, I’d say the website sits in the middle, at structured AI-assisted coding. It had detailed briefs, my own judgement at every stage gate, and testing on the site itself rather than a review of the code. Those checks were mine, made by hand, which is what the middle of the scale looks like. Google says the right level of checking “depends on the stakes”, decided task by task, and I don’t think agentic engineering would have got me a better website. Start from the outcome you’re after, and it tells you how much you need to build with and around the AI tool.

Two routes leave one circle labelled AI. The website route passes two checks, the writing route passes six, and both arrive ticked. Caption "only build what the outcome needs".

From Reimagining to Deploying a Writing Process

Building a pipeline for my writing was a different challenge altogether. I’m dyslexic, though I only found out late in life, and writing felt like a chore for as long as I can remember. Earlier in my career, I spent a lot of time writing reports and RFPs in sales environments. It was never my favourite thing, and it took me a long time. Later on, as tools like Grammarly and autocorrect came in, I probably used all of them, but none of them gave me the confidence to really want to write.

In the past and early in my career, some of that was imposter syndrome. Why would anyone want to read what I thought? When it came to blogs, the question was whether the return on my time was even worth it. I’d nearly say that without this AI-enabled pipeline I’ve built, I still wouldn’t be writing.

I’d used apps like Hemingway in the past to get a sense of how clear my writing was, and I can say today I still don’t know what an adverb is. So there was a lot of trial and error in building a pipeline that could take me from brainstorming an idea to a final published piece. I had to work out what the stages would be, what outcomes I wanted, where the quality bars sat, and whether a draft needed to go back through a step before it moved on. I certainly like to learn by doing, so even building the system was something I enjoyed.

Writing Has No Test, So I Built My Own

Writing is more personal and individual than code. Where code has tests that tell you it works, writing has no test that tells you it’s right, so you have to build your own. Each piece published moves through the same steps, and I’m involved at every stage:

  • Collecting the Material: It starts in Obsidian, with my working notes, my voice notes and clippings I’ve taken from third-party content, with my own take written into each clipping. The voice note I created for my previous piece, for example, ran to just over 8,000 words of me talking it through.
  • Agreeing the Idea: That material is pulled together into one idea for a piece, through an interview where Claude asks the questions and tests what I actually think, and a research pass that finds what supports it and what doesn’t. Nothing moves on until I’ve agreed it.
  • Shaping the Outline: The agreed idea is shaped into a full outline, section by section, that I refine and only sign off when I’m happy.
  • Drafting: Claude writes the draft from the outline and from my own words, and it’s then checked against my voice files and given up to three quality reads by Codex before I see it. Even with that, there are multiple draft versions created, edited and refined.
  • Reviewing: Once a piece reaches review, each round brings a fresh set of AI reads, from Codex and from Claude. Cold reads come from readers who know nothing about the piece, and quality reads judge it against my standards. Then a check that every source is properly attributed, and a read against Lisa’s file. I read every version and every issue raised, again, and decide what changes I make.
  • Making the Images: A separate image pipeline runs alongside, from the brief to the finished image. Image design briefs are written, and Codex and I go through greyscale ideation rounds up to finished images based on my design system.
  • Publishing: Once it’s through review, it’s formatted for the site and goes live on jalco.ai, with an audio version alongside it, created using ElevenLabs.

Seven circles in a gently curving row, labelled material, idea, outline, draft, review, images and publish, each ticked. One head-and-shoulders outline above the row has a thin line down to every circle. Caption "involved at every stage".

The writing was never about speed, and it certainly wasn’t faster at first. It was frustrating early on. What it has done, I believe, is help me frame my thoughts in a much more meaningful and coherent way, for me and I hope for others. I could build a fully automated process, and it would be a lot faster, but that was never the point.

At the top of the Writing page on jalco.ai, I say the pieces are written as I think them through, not after the fact. That’s what the pipeline is for. It helps me take an idea I have at the start, frame it better and bring it all together. The pieces are never finished either. Some have had updates since, and it’s why the images use scrappy lines, because they’re not supposed to be visually perfect. They’re evolving ideas, on subjects I’m trying to bring some clarity to, and I hope they help other people build positive momentum in what they’re trying to do.

With the website pipeline I trusted, but verified by checking and testing, the deployed visual experience. With my writing pipeline, nothing moves to the next stage or gets deployed until I’m happy with every word and every image. It’s my work, it’s public, and my name is on it, so that sign-off is never automated. For me, the stakes were in the words, not the code.

I feel my writing is better than it was, from a quality perspective, and the process machinery is making the Doing, the hands-on craft, a more enjoyable experience. Before, it was drudgery and it was painful. Now I’m happier and more excited doing it, and I find it really fulfilling.

What Google Counts as Agentic Engineering

Google’s whitepaper names four things agentic engineering needs in a coding scenario. With my writing pipeline, while it has some code, code is not the primary output, but it still meets each of Google’s criteria:

  • Specs and Memory Files: One rules file, 24 instruction files, one for each workspace and step, and the files on me.
  • Automated Tests: 34 scripts, of which about a dozen are checks a draft or post has to pass.
  • Gates: Before a post goes live, a check stops it if an image is missing or the layout would break the audio. At the end of every working session, a close-out won’t let the session finish until the work the AI says it did has actually been verified.
  • AI Judges: Seven AI reads in every review round, five by Codex and two by Claude.

The Upkeep Is What Keeps It Working

The upkeep is the part any operator will recognise, because anyone who has run a process already does it. You write down what went wrong, and you fix it so it can’t happen again. Google’s whitepaper says the same thing about agents: “Add a rule every time the agent does something it should not do again.” What I learnt is that a rule only holds once it becomes a check.

So far 120 defects have been logged; 93 are closed and 27 are still open. Open is normal; it’s a working system.

The drafts got cleaner once they were checked before I saw them. Across my last three pieces, the drafts each one needed fell from 14 to 6 to 3, and the lines I changed in review fell from 1,488 to 747 to 440. Three pieces is a short run, and the outlines got more thorough over the same stretch, so the checks aren’t the only cause. The direction is still clear.

The bigger lesson is one an engineer will recognise. The four fixes I wrote as instructions didn’t solve the problems I was encountering. Four other fixes, built as automatic checks, did. One of those checks replaced a written rule that told the AI to read my voice files before drafting, and the rule was being skipped. Eight fixes is a short run, but the pattern hasn’t broken. A written rule relies on someone remembering it, and a check doesn’t.

Two rows of four dots joined by arrows. The top row, labelled rule, is ticked at the first dot and crossed at the next three. The bottom row, labelled check, has a ticked circle at every dot. Caption "a check doesn't rely on someone remembering".

Not every fix is a check. Two others changed how the work is done. My voice faded in long drafts, because one file was trying to describe how I write everywhere. It was split by context, so published pieces, social posts, business documents, etc. each have their own voice. Another draft passed every rule I’d written and still didn’t sound like me, because it had been written from the outline, in the outline’s language, rather than from what I’d actually said. Now my own words are pulled out of my notes before drafting starts, and the draft is written from those.

Google’s closing line says it better than I would: “Generation is solved. Verification, judgment, and direction are the new craft.”

AI Is a Tool, People Provide Judgement and Direction

So where do I rank? I certainly wouldn’t have called myself an agentic engineer, but when I looked at what I’d built against Google’s scale, my writing pipeline was working at that agentic engineering level. The website sat in the middle, which is where it needed to be. If you think in systems or processes, you can do this too.

The person building the harness can be anyone in an organisation, at any level, including people who aren’t technical and aren’t steeped in this language. Put engineering in the name, though, and a lot of those people will decide it’s something somebody else does, in one of two ways. Either they don’t feel they can do it, or they feel there are other people in the organisation who can do it for them, and better. Neither is right. Defining the work by job function can cost people the curiosity to dive in, figure it out, fail, learn, test different things and just try.

Two identical circles. A dashed path runs into the first, labelled the work, and ends at a tick. A dashed path towards the second, labelled engineering, stops short and splits, one branch looping back, labelled can't do it, the other turning away, labelled someone else can. Caption "engineering in the name can be a roadblock".

To me, calling this process work “engineering” creates a stack ranking of experience and expectations, and for anyone without high agency, that can become a roadblock. People without that agency could be a lot of your organisation, and enabling those people build positive momentum is something I wrote about in AI Makes the Work Faster. Enabling People Makes It Compound.

AI is still a tool. It still needs a human in the loop, and the rule is still trust, but verify. I don’t anthropomorphise my tools; they’re there to help me. There’s a lot of noise about AI coming to take people’s jobs, but that’s not how I see these tools.

For anybody starting, do a bit of reimagining and research first, as I set out in Reimagine to Deploy. Ask yourself questions like, why would you build this? Is it going to save you time, and if so, what would you then do with that time? Where the freed up time goes when the build is cheap is the subject of Meaningful Doesn’t Need the Word Better.

I can’t say I know exactly where every organisation would benefit. You only find that out by getting inside the work and understanding how people actually do it. Figuring out how those people can generate more meaningful outcomes is job number one.

I also understand that I wasn’t working to a set timeline, and most people won’t have that luxury. Nobody should be expected to figure all this out in their spare time on top of the work they already do, because it’s probably a multi-function initiative. People need to be enabled to do it, and to have someone doing it with them while they keep everything else going. That’s the part I like to help with.

We’re all knowledge workers, whether you’re an accountant, a marketer, a salesperson or an engineer. The word engineering doesn’t decide who gets to build with these tools. If you can put a harness on the tool, you, the human, can win.

One plain head-and-shoulders outline joined by a line through a circle labelled AI, past a check, to a ticked target. Caption "you, the human, can win".

Sources

  1. Addy Osmani, Shubham Saboo and Sokratis Kartakis, The New SDLC With Vibe Coding, Google
    https://www.kaggle.com/whitepaper-the-new-SDLC-with-vibe-coding
  2. Jake Van Clief and David McDermott, Interpretable Context Methodology: Folder Structure as Agent Architecture
    https://arxiv.org/abs/2603.16021
  3. Design with Intent
    https://designwithintent.ai
  4. DHH, Omarchy
    https://omarchy.org
  5. John Leenane, Meaningful Doesn’t Need the Word Better
    https://jalco.ai/writing/meaningful-products
  6. John Leenane, AI Makes the Work Faster. Enabling People Makes It Compound.
    https://jalco.ai/writing/enabling-compounds
  7. John Leenane, Reimagine to Deploy
    https://jalco.ai/writing/reimagine-to-deploy
  8. John Leenane, Describe the Work, Not the Title
    https://jalco.ai/writing/describing-work
John Leenane

I'm John Leenane. I run JALCO, working day-to-day alongside teams across ops, finance, product and commercial strategy, helping them scale and expand into what's next. If you think I can help, let's talk.

Grab 30 mins