
This is the fifth agent in the Season 2 run of my AI Agent Series, where I break down every agent I actually run inside Purposeful Profits and across the brands I work with. This one changes what you believe about your own product, which makes it the most uncomfortable and the most valuable.
Picture starting a quarter knowing exactly what your customers say about you when you are not in the room. Not a sentiment score. Not a five-star average. The actual sentences, sorted by how often they repeat, with the six things that stop people buying listed in order of frequency. That is a different position to plan from than a gut feel and a team meeting.
What this used to cost, and why almost nobody does it properly
The advice has been the same for a decade. Go and read your reviews. Read your competitors' reviews. Copy the phrases that repeat. It is good advice and it is why most brands do it once, badly, and never again. A single analyst caps out at roughly 200 to 300 reviews a day, and that is reading them, not thinking about them. Three competitors across four platforms is a week of work before anyone has written a line of copy. So it gets scoped down to twenty five-star reviews and twenty one-star reviews on one site, which is enough to feel productive and not nearly enough to see a pattern.
Meanwhile the volume keeps growing. Your customers are talking on Trustpilot, on Google, in Reddit threads you have never seen, in Amazon reviews on a listing you may not even control, in app store reviews, under your TikToks, and in the support inbox. Something like 80% of that never gets analysed by anyone. The most valuable material is the unsolicited kind, which is exactly the material your own review widget cannot see, because you only collect from customers who bought and stayed subscribed.
Your review platform shows you the most flattering slice of your audience: people who bought, were happy enough to stay on the list, and were asked nicely. Everything that actually explains your conversion rate lives somewhere else.
What the agent actually does
It takes a brand name, a Shopify URL, a handful of Amazon ASINs and up to three competitors. Then it runs the following, mostly in parallel, while nobody watches it.
01
Build the source map
Before anything gets scraped the agent works out where the conversation actually is. Trustpilot and Google profiles, the Amazon listings and their variants, the subreddits where the category gets discussed, app store entries if there is an app, YouTube reviews, and the two or three niche forums that dominate the category. For supplements that might be a fitness forum; for skincare it is usually Reddit and a handful of review aggregators. The map is different for every brand and getting it wrong is the main reason DIY attempts return thin results.
02
Deploy scrapers in parallel
A team of scraper agents runs against every source at once rather than one after another. This is the step that turns a week into an afternoon. Each scraper returns raw records: the text, the rating, the date, the platform, and whether the reviewer was verified. Nothing is interpreted yet. The job here is coverage, and the hard floor is 200 deduplicated reviews before analysis is allowed to begin. In practice a real brand run returns thousands.
03
Deduplicate and quality filter
The same review often appears on four sites. Incentivised reviews read differently from organic ones and need to be weighted differently. Two-word reviews carry no signal. Delivery complaints are a logistics problem wearing a product complaint's clothes, and mixing them corrupts the whole analysis. This unglamorous step is where most homemade versions fall over, because a deduplicated set of 900 honest reviews beats a raw pile of 4,000 every time.
04
Cluster the themes and count them
Every review gets broken into the specific claims it makes, and those claims get clustered into recurring themes. Then, critically, they get counted. The difference between an objection that appears in 3% of reviews and one that appears in 31% is the difference between a footnote and a product roadmap item. Ranking by frequency is what turns reading into evidence.
05
Preserve the exact language
For every theme the agent keeps the verbatim quotes rather than paraphrasing them. This is deliberate. The value of voice of customer work is not the insight that people find your packaging hard to open, it is the sentence a real person wrote about wrestling with it in their kitchen at seven in the morning. That sentence is your next ad. Paraphrasing destroys the only thing you came for.
06
Produce the perception audit
The output is one document: the recurring positives in order of frequency, the recurring objections in order of frequency, the language customers use to describe the benefit, how the brand is described relative to each competitor, and the gaps where a competitor is being praised for something you also do but never say. It ends with the three things worth changing first.
The context layer is what stops it producing a sentiment score
Any language model can tell you a review is negative. That is not useful and it is why most feedback tools get opened twice and then ignored. What makes this agent worth running is the brand memory it carries in before it starts. It knows the product range and which SKUs are which, so it can tell a formulation complaint apart from a variant people simply do not like. It knows the margin structure, so a recurring request for a cheaper size can be flagged as a real opportunity or dismissed as one that would destroy contribution margin. It knows which objections the brand has already fixed, so an eighteen-month-old complaint about a discontinued formula does not get presented as a live problem.
It also knows what the brand is trying to become. A theme that reads as a weakness for a mass-market brand can be the exact thing a premium brand should lean into. Strong scent, thick texture, an acquired taste: those are objections in one positioning and proof points in another. The agent does not decide that. It reports the frequency and flags the tension, and the operator makes the call. That is the line between an agent that is useful and one that is confidently wrong.
What the output actually looks like
One document, usually eight to twelve pages, and it is not a dashboard. The first section is a ranked list of what customers love, with a count and three verbatim quotes under each. The second is the same treatment for objections. The third is the language section, which is the part I use most: a list of the phrases customers actually type, grouped by the job they are describing.
A real one from a wellness brand read roughly like this. Top positive theme, appearing in just over a third of reviews: it worked faster than they expected. Second: it did not upset their stomach, which was framed against a named competitor. Top objection, in around a quarter of reviews: the subscription was hard to pause. Second objection: the scent. The team had never discussed the scent. Their ad account had eleven active creatives and every one of them led with ingredient quality, which appeared in under 5% of what customers wrote.
That is the shape of the change. Before, the brand argued about positioning in a meeting. Now the argument has a frequency count attached to each side, and the losing position loses in about four minutes.
Where it still needs a human
An agent described as flawless is an agent nobody should trust, so here are the limits. Public reviews skew to the extremes: people write when they are delighted and when they are furious, and the quiet middle where most of your revenue sits is under-represented. The agent notes this rather than pretending the sample is clean, but it means a perception audit tells you about the edges of your customer base with much more confidence than the centre. Pair it with post-purchase survey data if you want the middle.
It also cannot tell you what people who did not buy think, which is the most expensive missing dataset in DTC. A one-star review is a customer who at least got far enough to be disappointed. The person who bounced off your product page at eleven seconds leaves no text anywhere. And it cannot tell you whether a frequent objection is worth fixing. That is a margin and roadmap decision that belongs to the operator, informed by the count rather than replaced by it.
Inside the system
How we build this for brands
The mining agent is the front of a chain rather than a standalone report. What it produces feeds straight into the VOC engine that turns customer language into positioning and TikTok and Meta ad creative, so the phrases with the highest frequency count become the hooks running in the ad account within days. The same document feeds the lifecycle flows built and deployed in Klaviyo, which is why the win-back email answers the objection that actually appears most often rather than the one someone guessed at.
Further down the chain the same language shapes the outreach and relationship agents, which are trained on the brand voice and thousands of real customer touch points, and it informs who gets invited to the small IRL experiences we run: pop-up tastings, run clubs, sauna sessions, where the right hundred customers matter more than the wrong thousand. Part of this runs live for portfolio brands today; the full system is what we deploy when we take a brand on.
Build Your Agent Team
Get your own AI agent team working on your brand
Caner can help you build your own AI agent teams and unlock growth for your brand. This is not a template or a SaaS tool. It is a custom system designed around your data, your workflows, and your growth targets.
Book A DemoFrequently asked questions
What is a VOC mining agent and what does it do for a DTC brand?
A VOC mining agent is an automated system that collects customer commentary about your brand from every public channel it appears on, then analyses it at scale. It runs parallel scrapers across Google reviews, Reddit threads, Trustpilot, Amazon listings, the app stores, YouTube comments and niche forums, deduplicates what it finds, and produces a brand perception audit: the recurring positives, the recurring objections, the exact phrases customers use, and how your brand is described relative to competitors. The output is not a sentiment score. It is a document you can write ads, product pages and email flows from.
Can I build a VOC mining agent myself?
You can build a version of it. Scraping a Trustpilot profile or an Amazon listing is a solved problem and there are off-the-shelf actors that do it well. The hard part is everything downstream. Deduplicating the same review syndicated across four sites, telling an incentivised review apart from an organic one, separating a product complaint from a delivery complaint, and knowing which of six recurring objections actually blocks purchase rather than just annoying people after the fact. That judgement layer is where the value sits, and it is the part that takes months to encode rather than a weekend.
How many reviews do you need before the analysis is trustworthy?
The floor we work to is 200 deduplicated reviews, and we usually target thousands. Below 200 you are reading noise and mistaking loud individuals for patterns. Between 200 and 500 the recurring themes stabilise and you can trust the top three or four. Above 1,000 you start to see the smaller, more useful signals: the specific objection that only appears in one customer segment, or the use case nobody in the business knew existed. Most brands have far more public commentary available than they think, because they only ever look at their own review widget.
How long does it take to set up VOC mining for an ecommerce brand?
The first full run takes a working day, most of which is the agent running rather than anyone doing anything. Setting up the sources properly takes longer: identifying every place your brand and your three closest competitors are discussed, confirming the right Amazon ASINs, and finding the Reddit threads that matter. Expect a week to get the source map right, then the run itself becomes something you repeat monthly or quarterly in a few hours of compute with no human input.
How is this different from the sentiment dashboard in my review platform?
Your review platform only sees the reviews it collected, which are the ones you asked for, from customers who completed a purchase and stayed subscribed to your email list. That is the most favourable slice of your audience. It cannot see the Reddit thread where someone talks your product down, the Amazon one-star that names a competitor, or the app store review from someone who never made it through checkout. A VOC mining agent goes to where the conversation actually happens rather than where you have permission to collect it.
What do you actually do with a brand perception audit once you have one?
Three things, in order. Rewrite the product page around the objections that appear most often, because every one of them is currently costing you conversions silently. Rewrite the top of the ad account using the exact phrases customers use to describe the benefit, because customer language outperforms brand language almost every time. Then feed the recurring positives into your review request and post-purchase flows so you are prompting for the proof you know converts. The audit is an input to copy, not a report to file.
About the author
Caner Veli built Liquiproof to global distribution across 3,000+ retailers, then exited. He now runs Purposeful Profits using a combination of operator strategy and AI-powered systems he has built and uses daily, having 10x'd monthly revenue in his own business in the last 90 days.