
The Monday workbook was never the expensive part. The expensive part was that it sat at the front of the week, ate the hours when my head was clearest, and left me doing the thinking about the numbers in the leftover ten minutes. Four hours of clerical work bought me ten minutes of judgement. That is the wrong ratio in any business, and it is a fatal one in a brand where the margin is thin enough that the judgement is the whole job.
I tried to solve it the normal ways first. A VA, which meant I now spent an hour reviewing someone else's version of the same numbers. A reporting tool, which gave me a dashboard nobody read because it answered questions I was not asking. Then I tried asking Claude in a chat window, which was worse, because it could tell me what to do with the numbers but could not go and get them. I still had to open six tabs, copy, paste, and check the formulas.
What fixed it was Claude Code. Not because it is smarter than the chat window, it is the same model underneath. Because it runs somewhere that can actually reach my data and do the work, and then hand me the finished thing.
The difference between a chat window and a runtime
A chat window is a very good advisor with no hands. You describe your situation, it describes what to do, and every single step of the doing lands back on you. The value it produces is capped by how fast you can execute on advice, which for most operators is the actual bottleneck. You do not have an ideas problem. You have a fourteen-open-tabs problem.
Claude Code runs on a machine. It has a filesystem, a terminal, and connections to whatever tools you have given it access to. That means it can read your actual Shopify export rather than a summary you pasted, write the file rather than tell you what to put in it, run the script, check its own output, and commit the change. The output of a session is not a paragraph of advice. It is a workbook, a deployed email flow, a set of ad variants, a reconciled ledger.
That distinction is why search demand for the agent SDK went from around 50 monthly searches in May 2025 to roughly 14,800 by April 2026. Operators worked out that the interesting question is not what the model knows. It is what the model is allowed to touch.
A chat window makes you a faster researcher. A runtime makes you a smaller team that ships more. Those are not the same upgrade, and only one of them shows up in your P&L.
One real workflow, start to finish
Abstract examples are useless here, so this is the actual Monday reporting job as it runs now, for a wellness brand in the portfolio. It fires at 7am before I am at a desk. Nothing in this list is aspirational.
It reads the brand context first
Before touching data, the agent reads a written context file for that brand: the KPI definitions we agreed, which SKUs are seasonal, the thresholds that count as an exception, and the fact that this brand counts subscription revenue separately. This file is the difference between an agent and a search box. It is roughly two pages, written by me, updated whenever something changes.
It pulls the numbers from source
Week-to-date revenue, orders and AOV from Shopify. Blended and channel-level spend, MER and new customer CAC from Triple Whale. Flow and campaign revenue, list growth and unsubscribe rate from Klaviyo. Not a screenshot, not a paste. Direct reads through connected tools, with the previous four weeks pulled alongside so every figure lands next to a trend rather than on its own.
It populates the live workbook and checks its own maths
The figures go into the existing workbook, in the existing structure, so the client opens the same file they always have. Then it re-derives the calculated fields independently and compares them against the workbook formulas. Where the two disagree, it flags rather than overwrites. That step exists because of the mistake in the note at the top of this page.
It writes the exceptions, not the summary
A recap of numbers I can already see is worthless. What I want is the four lines that say what broke: replenishment flow revenue down 22% week on week, one ad set consuming 31% of spend at half the account MER, list growth flat for the third week. Each flag names the metric, the size of the move, and the window it happened in.
It hands me the decision, and stops
The agent does not pause the ad set. It does not touch the flow. It produces the workbook and the exception list and stops there, because the next step is a judgement call about a brand I have a relationship with. Deciding where that line sits is the most important design choice in any of this, and it is one an operator has to make, not a tool.
Four hours became about eight minutes of reading. The workbook is more accurate than my version was, because it never gets bored on the third tab. And the ten minutes of judgement that used to be the leftovers are now the entire meeting.
The context file is the whole trick
Most people who try this and give up have the same problem. They ask for a good output without ever telling the agent what good means. The model has no idea that this brand's Q1 always dips because of a January detox spike the year before, or that the founder hates the word "journey", or that a 3% conversion rate is excellent in their category and mediocre in another.
So the first real asset you build is not a workflow. It is a written record of everything you know that has never left your head. Brand voice with banned words. Margin structure per SKU. Which channels are load-bearing. What last quarter's failed test proved. Which numbers are allowed to move 20% without anyone panicking. Every agent I run reads that file before it does anything else.
This is also the part that compounds. A workflow saves you the hours it was built for. A good context layer makes every future workflow start at 80% instead of zero, and it is the closest thing to institutional memory a small brand can build without hiring for it.
Three things it gets wrong, honestly
Anyone writing about a tool without naming its failures is selling something. These are the three that have actually cost me.
Context is a consumable, and it is expensive
Every file the agent reads and every tool response it receives stays in that session's context permanently. A long-running job quietly accumulates until it is burning far more than the work is worth, and if you run several agents in parallel each one carries its own full context. Anthropic changed the usage caps four times between March and June 2026, which tells you how unsettled this still is.
The practical fix is discipline about scope. One workflow per session. Pull the narrow slice of data the job needs rather than the whole export. Kill long-running sessions rather than letting them idle. If a workflow feels expensive, it is almost always reading more than it needs to.
It fails silently and confidently
This is the dangerous one. An agent that successfully pulls six days of data instead of seven does not error. It produces a clean, well-formatted, entirely confident report built on incomplete input, and nothing in the output looks wrong. That is how I ended up presenting a flat week that was actually a good one.
Every workflow that touches a number now carries a validation step that I designed: row counts, date-range checks, and an independent re-derivation of the calculated fields. If a workflow cannot be checked cheaply, it does not get to run unattended.
It has no taste
It will write a competent ad. A competent subject line. A competent landing page. Competent is exactly the problem, because competent does not stop a thumb in a crowded category and it does not build a brand anyone remembers.
Where it earns its place in creative work is volume and variation on a strong idea, not the idea itself. I still write the angle. I still pick the winner. The agent produces the fifteen executions of it that I would never have had the patience to make, and takes the first pass at the reading of what performed.
Where to start if you run a brand
Do not start with the most valuable workflow. Start with the most boring one you can describe precisely. The job you could hand a competent new hire with a two-page brief is the job an agent will do well. The job that requires you to be in the room is not, and trying to automate it first is how people conclude none of this works.
Write the context file before you write the workflow. Give the agent read access only, until you have watched it run enough times to know where it drifts. Run it alongside your manual process for two weeks and compare the two outputs every time. When the agent's version stops surprising you, drop the manual one.
Then do the next one. The reason this compounds is not that any single workflow is transformative. It is that the tenth one takes an afternoon instead of a fortnight, because the context, the connectors and the patterns are already there.
What this looks like in practice
The Monday workbook was the first one. It is now one of about a dozen that run across my own business and the brands I work with. Payment reconciliation against issued invoices. Review mining across Google, Reddit, Trustpilot and Amazon into a positioning brief. Inbox triage before I open a laptop. Klaviyo flows drafted and deployed. Ad creative generated from a brief. This article was researched, written and published by one of them.
The compound effect is not the hours, though the hours are real. It is that the work which used to get skipped now happens every week without negotiation. Nobody ever skipped strategy because it was unimportant. They skipped it because Monday was full.
In the last 90 days I have 10x'd monthly revenue in my own business. A meaningful part of that is simply that the highest-value work stopped competing for calendar space with work that a system should have been doing all along.
Inside the system
How we build this for brands
When we take a brand on, the runtime is the layer underneath everything else. We build the brand context file first, trained on the voice, the margin structure and thousands of real customer touch points, then connect the data the agents need: Shopify and ad platforms for the profit and cash-flow dashboard, Klaviyo for the lifecycle flows that get built and deployed rather than just recommended, and a VOC engine that turns reviews and support messages into positioning and TikTok and Meta creative. A reporting agent runs weekly on top of all of it and surfaces leakage or risk before it shows up in the bank.
The point is never the automation for its own sake. It is that the founder gets their week back and the strategic work stops being the thing that gets pushed. Part of this runs live for portfolio brands today; the full system is what we deploy when we take a brand on.
Growth Audit
Find Out Which Hours Your Brand Is Wasting
I will map the recurring work sitting on your team's calendar, identify which of it a system should be doing, and show you what the week looks like once it is gone. Then we talk about the growth work that finally has room to happen.
Book Your AuditFrequently asked questions
What is Claude Code and how is it different from using Claude in a chat window?
Claude Code is an agent runtime rather than a conversation interface. A chat window can only return text to you, so you remain the person who copies the answer somewhere useful. Claude Code runs on a machine with access to files, a terminal, and connected tools, so it can read your actual data, write the file, run the script, commit the change, and report back. The practical difference for a DTC operator is that the chat window produces advice and the runtime produces finished work.
Do I need to be a developer to use Claude Code for ecommerce operations?
You need to be comfortable in a terminal and willing to read what the agent did before you trust it, but you do not need to write code. Most of the work is describing the job precisely, connecting the right data sources, and building up a written context file that tells the agent who you are, what the brand sounds like, and what good output looks like. That is operator work. The engineering-adjacent part is setting up connectors safely, which is a one-time job.
What can Claude Code actually automate for a DTC brand?
The workflows that pay back fastest are repetitive, data-heavy and low-judgement. Weekly KPI reporting from Shopify, Triple Whale and Klaviyo into a live workbook. Payment reconciliation against issued invoices. Review mining across Google, Reddit, Trustpilot and Amazon into a positioning brief. Ad creative and caption production from a brief. Inbox triage and draft responses. Campaign tracking updates in Notion. Anything that requires taste, negotiation, or a relationship stays with you.
What are the real limitations of Claude Code for business operations?
Three matter. Context cost: every file read and tool response stacks in the session, so long jobs get expensive and the usage caps changed four times between March and June 2026. Silent partial failure: an agent that pulls 6 days of data instead of 7 will still hand you a confident report, so anything touching numbers needs a validation step you control. And taste: it will produce a competent ad, a competent email and a competent deck, and competent is not what wins in a crowded category.
How long does it take to get a Claude Code workflow running reliably?
A first version of a single workflow usually takes an afternoon. Getting it reliable enough to run unattended takes two to four weeks of running it alongside the manual process and fixing what breaks. The gap is almost never the model. It is the edge cases in your own data, the connector that rate-limits at the wrong moment, and the undocumented judgement in your head that the agent could not read. Write that judgement down and reliability improves quickly.
Is it safe to give an AI agent access to Shopify, Klaviyo and bank data?
It is safe if you scope it properly. Give read access by default and write access only where you have explicitly decided the agent should act. Keep destructive operations behind a human confirmation. Use separate API keys per workflow so you can revoke one without breaking everything. Never store credentials where the agent writes its output. The risk is not the model deciding to do something malicious. The risk is a badly scoped permission plus an ambiguous instruction.
About the author
Caner Veli built Liquiproof to global distribution across 3,000+ retailers, then exited. He now runs Purposeful Profits using a combination of operator strategy and AI-powered systems he has built and uses daily, having 10x'd monthly revenue in his own business in the last 90 days.