
The short version: Testing a platform yourself finds feature gaps. Testing it with real customers finds confusion, and confusion is what actually costs you orders. The two failure modes are completely different, which is why a platform that passed every solo test can still produce a bad first weekend. Run a soft launch with eight to twelve regulars, chosen because they will forgive a problem rather than because they spend the most, tell them plainly that they are helping, and watch for the things you cannot see from the inside: where they hesitate, what they message you about, and who quietly does not complete an order.
Because you already know the answers.
When you test your own storefront, you know what the products are called, where the cutoff is, which pickup point is which, and what "Saturday collection" means. Your customers know none of that, and every gap between what you meant and what they read is an order that takes longer or does not happen.
Solo testing reliably finds:
Solo testing reliably misses:
That second list is why the live test exists. It is not a repeat of the first test with more people; it is a different test.
Eight to twelve people, chosen for forgiveness rather than value.
The instinct is to test with your best customers, and it is the wrong instinct. Your highest-spending customers are the ones you least want to give a confusing experience to, and they are often the least tolerant of friction because they are used to your existing process working.
Choose instead:
That third and fourth are the ones people skip. A test group of easy orders tells you the easy path works, which you already knew.
The truth, briefly, framed as a favour rather than an apology.
Something like:
> "I'm trying a new ordering page and would love a few people to use it this week and tell me if anything's confusing. Same bread, same pickup, just a different link."
That does four things: it explains why, it sets the expectation that feedback is wanted, it reassures on the things that have not changed, and it does not apologise. An apology tells people to expect a problem.
What not to do:
Six signals, and only two of them are numbers.
Signal three is the most useful and the easiest to record. Keep a list. By the end of the week you will know exactly which three sentences are missing from your listings, and it will not be the three you would have guessed.
Our guide to writing product descriptions that sell food online covers fixing them once you know, which is a much faster job than guessing at it in advance.
One complete order cycle, whatever that means for you.
For a weekly baker, that is one week: orders open, cutoff passes, collection happens, everyone goes home. For a monthly drop, one drop. The cycle is the unit because most of what you are testing happens at the ends rather than in the middle.
What you specifically want to observe:
One cycle is enough. Two is better if the first was unusual. More than that is delay rather than testing, and every extra week is a week of running two systems.
Fix it for that customer immediately, then decide whether it is a platform problem or a setup problem.
The immediate response is always the same: take the order the old way, apologise once, move on. A test customer who cannot complete an order should not lose their bread over it.
Then classify what happened:
That last distinction matters because vendors routinely blame platforms for the third kind. If two people ordered for the wrong day, the platform is probably fine and your day labels are not.
Have one before you start, because you will feel calmer and it will almost certainly be unused.
That fourth point is the one that keeps this low-risk. A soft launch with a dozen people can go badly and cost you nothing. A public announcement followed by a bad first weekend costs you goodwill you have to rebuild.
When the same cycle runs twice without a new problem.
The sequence that works:
Step four is the honest one. Two cycles of the same problem is a signal, and the correct response is sometimes to stop rather than to keep adjusting. A platform that needs three rounds of workarounds to handle an ordinary week will need them every week afterwards.
On step three, the announcement is four messages rather than one: a heads-up before the date, the announcement on the date, a reminder a week later, and an individual reply with the link to anyone still using the old channel. Our guide to moving a food business off Instagram DMs in one weekend covers that sequence in its most common form.
Worth a moment, because "testing on people" deserves a little care even at this scale.
The standards that make it fine:
That third point has actual obligations attached. A trial platform holding customer names, addresses, and payment details is subject to the same expectations as a live one, and the FTC's privacy and data security guidance for businesses sets out what those involve. If you abandon the platform afterwards, delete the test data rather than leaving it in a system you no longer use.
The IRS's recordkeeping guidance is the other half: orders taken during a test are real sales and belong in your records regardless of which system processed them.
Things you could not have learned any other way.
By the end you should know:
That last one is worth taking seriously. A platform that technically does everything and that you dread using every Thursday is a bad choice, and one cycle is enough to know.
Act on it within the same week, while the detail is still specific.
Point three is the one that pays back beyond this project. A customer who was asked for their opinion and then told what changed because of it is measurably more loyal than one who simply had a good experience. Our guide to getting repeat customers covers why, and a soft launch is one of the few natural opportunities to do it sincerely.
If you are still choosing which platform to put in front of those twelve people, set up a trial with your real catalog first and run your own eight tests before involving anyone else. A live test should find confusion, not configuration mistakes you could have caught alone.
If you want to run this against a platform built around collection days and per-location cutoffs, Homegrown is $10 a month billed annually with 0% commission and 2.9% plus $0.30 processing published up front, with a 7-day free trial. It handles pickup at each place you sell with its own schedule and cutoff, local delivery with a radius and a route, and sales tax calculated, filed, and remitted in all 50 states. The honest bounds: no point-of-sale, no national shipping, no drop countdowns, and no app ecosystem. Seven days covers one weekly cycle, which is exactly the unit this article recommends, so starting it the week before a normal order week with a dozen regulars is the test that produces evidence rather than impressions.
Every platform below shows the same four commercial facts, because a table that lists one platform's transaction fee and not another's is not a comparison. "Not published" means exactly that: the company does not state it publicly.
| Platform | How long you get to test | Subscription (annual) | Free trial | Platform fee | Card processing |
|---|---|---|---|---|---|
| Bakesy | 30-day trial, the longest here | $9.99/mo Standard (monthly only) | 30-day free trial | $0 platform fee | 3.9% + $0.30 processing, optional |
| Bake.Shop | 14 days with full features | $149/yr (= $12.42/mo) | 14-day free trial | $0 platform fee (0% commission) | 2.9% + $0.30 processing |
| Homegrown | 7 days | $10/mo billed annually | 7-day free trial | $0 platform fee (0% commission) | 2.9% + $0.30 processing |
| Big Cartel | 7 days on Platinum | Platinum $12/mo ($144/yr) | 7-day free trial | $0 platform fee | Your own provider, so 2.9% + $0.30 typical |
| Barn2Door | No trial, demo only | $119/mo + $399 one-time setup | No free trial, demo only | $0 commission | 2.9% + $0.30 processing |
| Cottage CMS | Free tier, so test indefinitely | Free tier; Pro $200/yr | No trial needed, free tier | Platform fee not stated | Square's processing rate, not restated |
Because solo testing finds feature gaps and live testing finds confusion, and confusion is what costs orders. Nothing is ambiguous to its author, so you cannot see the unclear parts of your own storefront.
Eight to twelve. Enough to surface patterns, few enough to pay attention to each one, and small enough that a bad week costs you nothing you cannot repair personally.
Regulars who will forgive a problem, including one who finds technology difficult, one who orders something complicated, and one who normally orders by text. Not your highest spenders, and nobody with a time-critical order.
Yes, briefly, framed as a favour rather than an apology. If something is confusing and they do not know it is new, they conclude your business has got worse rather than that a page needs fixing.
One complete order cycle: orders open, cutoff passes, collection happens. Two if the first was unusual. More than that is delay, and every extra week means running two systems.
Take that order the old way immediately, then classify what happened: a setup problem you fix, a platform problem you score, or a communication problem where the page did not say something. The third is by far the most common.
After two cycles run without a new problem. A soft launch that goes badly costs nothing; a public announcement followed by a bad weekend costs goodwill you then have to rebuild.
Solo testing and live testing find different things. The first finds what the software cannot do; the second finds confusion, and confusion is what actually loses orders. A platform can pass every test you ran alone and still produce a bad first weekend.
So run a soft launch: eight to twelve regulars, chosen for forgiveness rather than spend, told plainly that they are helping, across one complete order cycle. Keep the old channel open, so nothing is at risk, and do not announce anything publicly until a second cycle runs clean.
Then watch the signal nobody records: the questions. Every question is a sentence missing from your page, and by the end of the week you will know exactly which three are missing. It will not be the three you would have guessed, which is the entire reason for doing this with real people.
