The most expensive thing a DTC brand does is guess. Guess which objection is killing the product page. Guess which benefit to lead the ad with. Guess whether the returns are a sizing problem or a photography problem. Every one of those guesses costs a creative cycle, and a creative cycle on paid social is two weeks and a few thousand pounds before you learn anything.
The answers already exist. They are sitting in your Amazon review section, in the Trustpilot page you have not opened since the bad month, in a Reddit thread where three people compared you to a competitor in language no copywriter would ever have written. The problem was never that the data was missing. It was that reading four thousand reviews by hand is not a job anyone will actually do, so nobody does it, and the brand goes on guessing.
The job first: finding the words your customers already use
I have never met a brand that describes its product the way its customers do. Founders describe mechanism. Customers describe consequence. The founder says clinically proven ceramide complex. The customer says my skin stopped stinging when I put makeup on. Those two sentences are about the same product and only one of them stops a thumb.
This is why voice-of-customer work outperforms brainstorming so consistently. Headlines sourced from real customer language beat marketer-written headlines by roughly 19%, and rewriting a page in that language has been shown to move conversion by multiples rather than percentages. The gap is not talent. It is that one version was invented in a meeting and the other was retrieved from people who paid money.
You are not trying to write better copy. You are trying to stop writing copy and start transcribing it.
Why Apify, and what it actually is
Apify is a cloud platform that runs pre-built scrapers against public web pages and hands you back structured data. Each scraper is called an Actor. You point one at a URL or a search term, set a few inputs, and it returns a dataset you can pull into a spreadsheet or feed straight into an AI agent.
I chose it for one specific reason: coverage in one place. Before Apify I was running a Chrome extension for Amazon, a paid tool for Trustpilot, a manual export for Shopify reviews, and copy-pasting Reddit threads into a document. Four tools, four formats, four exports that had to be reconciled by hand. Apify collapsed that into a single interface with a single output shape, which is what made the work automatable rather than merely possible.
The part that mattered more than the scraping was the API. Because every Actor can be called programmatically and every run returns a dataset by ID, an AI agent can fire ten scrapers in parallel, wait for them, deduplicate the results, and analyse them without a human touching a CSV. That is the difference between a research tool and an operating system. The scraping is the boring half.
One real workflow, start to finish
This is the exact sequence I run before writing a single line of creative for a brand. It takes under half an hour of my attention. Here is what happens in it.
Define the question, not the dataset
Before anything runs I write down the one thing I need to know. Usually it is a version of: why do people who nearly buy this decide not to. That single sentence determines which sources matter. Skip it and you end up with eight thousand rows and no conclusion, which is the most common way this work fails.
Fire scrapers across every channel at once
Amazon for the product-level defects and the sizing complaints. Trustpilot for service and delivery. Reddit for the honest competitor comparison nobody leaves on a brand's own site. The app stores if there is an app. Shopify review apps on the brand's own store and on two competitors. These run in parallel, not in sequence, because waiting on each one in turn is what makes people give up.
Deduplicate and set a floor
Syndicated reviews appear in multiple places and will skew everything if you let them. I dedupe on review text and author, then check the count. Below two hundred deduplicated reviews for a single product I do not draw conclusions, I widen the sources. Volume is what turns an anecdote into a pattern.
Cluster into objections and triggers
Every review is sorted into one of two buckets that matter: what nearly stopped them buying, and what made them buy anyway. Everything else is noise. The output is a ranked list of recurring objections with a frequency count and, critically, the verbatim phrasing customers used for each one.
Write hooks against the objections
The top three objections become the hooks. Not the top three benefits, the top three objections, because the objection is what the reader is already thinking and naming it is what earns the next three seconds. The verbatim phrasing goes in unedited wherever it survives a read-aloud test.
Ship variants and check the language landed
Five ad concepts, three formats each, into test. The metric I watch first is not ROAS, it is hook rate, because that tells me whether the customer language actually resonated or whether I picked the wrong objection out of the list.
Done manually, steps two through four are a week of someone's life and they will do it badly, because reading two thousand reviews without a system means remembering the vivid ones rather than counting the common ones. Automated, it is a run I kick off and come back to. The judgement call in step five is mine and always will be. Everything before it is retrieval, and retrieval is exactly the kind of work that should never touch a founder's calendar.
Where it breaks, honestly
Actor quality is inconsistent and this is the thing nobody warns you about. Most Actors are built and maintained by independent developers, so two scrapers aimed at the same site can return meaningfully different completeness, and a well-reviewed Actor can quietly degrade when the target site changes its markup. Worse, several review Actors cap results per company without making it obvious, so you finish a run convinced you pulled everything when you pulled the first few hundred. I now always sanity-check the returned count against the review count visible on the page before I trust a dataset.
The cost model is the second trap. You pay across compute and proxy bandwidth rather than a flat per-run fee, and on Amazon residential proxies are effectively mandatory because datacentre addresses increasingly return listings with missing prices or the wrong marketplace. A badly scoped run that crawls far more pages than you intended does not fail, it just costs more than you planned. Scope the input tightly, run a small test first, then widen. I have burned credit on a run that returned nothing useful because I pointed it at a category instead of a product.
And the honest limitation that has nothing to do with the tool: scraping gives you volume, not truth. Reviews over-represent the delighted and the furious, and they under-represent the largest group of all, the people who considered you and silently left. Apify cannot see those people. That is why I still run post-purchase surveys and still get on calls, and why I treat the scraped dataset as the map rather than the territory. It tells you where to dig. It does not tell you what is buried.
What this looks like in practice
A supplement brand I work with had flat conversion on their hero product and a founder convinced the price was the issue. They had run three discount campaigns in a quarter, each one lifting volume and shredding margin, none of them changing the underlying rate.
We pulled reviews across their own store, two competitors, and Amazon. Price appeared as an objection in a small minority of them. The dominant recurring theme was uncertainty about how long the product takes to work, expressed over and over in slightly different words, mostly by people giving four stars and hedging. Nobody had ever put a timeline on the product page because internally everyone knew the answer and assumed it was obvious.
We added the timeline to the page, rebuilt the top-of-funnel hooks around it in the customers' own phrasing, and stopped discounting. Conversion moved, and it moved without giving margin away, which is the only kind of conversion lift worth having. The insight cost a scraping run. The eight months of discounting before it cost considerably more.
Inside the system
How we build this for brands
When we take a brand on, the VOC engine is usually the first thing we stand up, because almost everything downstream depends on knowing how customers actually talk. It mines reviews and support messages across every channel the brand appears on, deduplicates them, and turns the recurring objections and triggers into positioning and into TikTok and Meta ad creative. It reruns on a cadence rather than once, because customer language shifts and a perception audit from March is not the one you want briefing October's creative.
The same underlying data then feeds the lifecycle flows we build and deploy in Klaviyo, so the win-back email argues against the objection that actually loses customers rather than a generic one, and it feeds the creator and buyer discovery agents so outreach is written in language the audience recognises. Part of this runs live for portfolio brands today; the full system is what we deploy when we take a brand on.
Growth Audit
Find Out What Your Customers Are Actually Objecting To
I will mine your reviews and your competitors' reviews at scale, rank the objections that are costing you conversion, and show you the exact language to answer them with. Then we build the creative and the flows against it.
Book Your AuditFrequently asked questions
What is Apify and how do DTC brands use it?
Apify is a cloud platform that runs pre-built scrapers, called Actors, against public web pages and returns structured data. DTC brands use it mainly for three jobs: pulling customer reviews from Amazon, Trustpilot, Shopify review apps and the app stores to mine customer language for copy; monitoring competitor catalogues and prices; and building creator or affiliate lists from social platforms. You do not need an engineer to run it, but you do need to know what question you are answering before you start.
Why does scraping customer reviews improve ad copy?
Because customers describe your product differently to how you describe it. Headlines sourced from review mining or customer interviews outperform marketer-written headlines by around 19%, and rewriting a landing page in voice-of-customer language has been shown to lift conversion by two to five times. Reviews give you the exact objection, the exact trigger, and the exact phrasing at volume, which is something a brainstorm cannot produce.
How many reviews do you need for voice of customer research to be reliable?
Two hundred deduplicated reviews is the floor for a single product. Below that you are reading anecdotes and you will over-index on whichever complaint you happened to read last. For a full brand perception audit across a category, aim for the low thousands spread across at least four sources, because every platform has its own bias. Amazon skews to product defects, Trustpilot skews to delivery and service, Reddit gives you the unfiltered comparison against competitors.
Is scraping reviews with Apify legal?
Collecting publicly visible information is generally treated differently to collecting personal or gated data, but platform terms of service and your own data-protection obligations still apply, and this is not legal advice. In practice the rules I hold to are simple: only public pages, never anything behind a login, never personal data beyond the display name attached to a public review, and never republishing scraped review text verbatim as if it were your own marketing claim. Use it to understand language, not to lift content.
What are the main limitations of Apify?
Three real ones. Actor quality varies widely because most are built by independent developers, so two scrapers pointed at the same site can return very different completeness. Costs are usage-based across compute and proxy bandwidth, which means a badly scoped run can quietly cost far more than you expected, and residential proxies are effectively mandatory on Amazon. And some Actors silently cap results per company, so you can finish a run believing you pulled everything when you pulled the first few hundred.
Can I use scraped reviews to write Meta and TikTok ad copy?
Yes, and it is the highest-return use of the data. The workflow is to cluster reviews into recurring objections and recurring triggers, pull the verbatim phrasing customers use for each, then write hooks against the objection and body copy against the trigger. You are not quoting reviews as testimonials, you are borrowing the vocabulary. The ads that win are almost always the ones that sound like the review section rather than the brand book.
About the author
Caner Veli built Liquiproof to global distribution across 3,000+ retailers, then exited. He now runs Purposeful Profits using a combination of operator strategy and AI-powered systems he has built and uses daily, having 10x'd monthly revenue in his own business in the last 90 days.
