Voice commerce is not a replacement for the online store. It is an input method that can make a small number of customer tasks faster: repeat an exact order, check a status, add a known item or begin a support request. It fails when the customer must compare, inspect or understand too much information by sound alone.
The promise of “shopping by conversation” has circulated for years. Meanwhile, the platform landscape has changed. Google closed custom Conversational Actions for Assistant in 2023, while Amazon continues to support defined Alexa shopping actions inside its own ecosystem. That history is a warning for independent businesses: do not build a commerce strategy around a platform feature you do not control.
A better approach begins with the customer task. Determine when speaking is genuinely easier than touching or typing, design confirmation and recovery around predictable recognition errors, and keep a visual or human alternative. Voice can then improve access and convenience without becoming a theatrical technology project.

Separate voice input from voice-only commerce
A customer can use speech in several ways. They can dictate into a normal search or form field, control an accessible website through speech-recognition software, speak to a business chatbot, use a phone assistant to trigger an app action, or interact with a smart speaker. These routes have different technical, commercial and privacy implications.
Voice input adds another way to operate an existing journey. Voice-only commerce tries to deliver the journey through sound without a screen. The first is broadly useful. The second should be reserved for tasks with little ambiguity and low decision complexity.
W3C accessibility guidance notes that speech recognition supports people with physical disabilities, repetitive-strain injuries and cognitive or learning disabilities, as well as people with temporary limitations. It also explains that properly coded controls, keyboard accessibility and matching visible labels help speech users operate websites. A business can therefore gain much of voice’s accessibility value by improving its ordinary website rather than creating a branded assistant.
Tasks most suitable for a voice layer
This is a task-design hierarchy, not usage data. Risk, user capability, language and environmental noise can change the appropriate channel.
Choose a task with low ambiguity
Voice performs best when the customer already knows the object and desired action. “Reorder the same 500-gram coffee beans as last month” contains a reference the system can resolve. “Find me a good coffee” requires preference discovery, product comparison and explanation.
Evaluate five dimensions:
- Object certainty: is the exact product, appointment or order identifiable?
- Choice count: can the options be understood without remembering a long spoken list?
- Consequence: what happens if recognition or intent is wrong?
- Reversibility: can the customer correct or cancel easily?
- Context: is the customer likely to be in a place where speaking and hearing are practical?
A spoken interface is weak for visually differentiated products, detailed specifications, unfamiliar prices, contracts and sensitive information. It can be strong for status, inventory, routine replenishment and simple scheduling.
| Customer task | Voice role | Required safeguard | Better fallback |
|---|---|---|---|
| Check delivery status | Read one current state and next event | Account verification and no unnecessary detail aloud | Link to visual tracking |
| Repeat a previous order | Resolve exact item and quantity from history | Price, item, quantity, address and final confirmation | Reviewable basket |
| Add a known item | Capture model or product identifier | Compatibility and duplicate checks | Visual product card |
| Book a standard appointment | Offer a few available times | Timezone, duration, location and cancellation terms | Calendar picker |
| Compare insurance or complex equipment | Gather initial needs only | No hidden recommendation or consequential decision | Structured comparison and adviser |
Design the conversation around repair
People pause, restart, change their minds and refer to objects indirectly. Speech recognition can confuse product names, accents, numbers and noisy environments. The central design task is therefore not the happy path. It is repair.
Confirm the part that carries risk. Repeating every word makes conversation exhausting; failing to confirm quantity, price or address creates real harm. Use concise prompts: “I found your previous order: two refill packs at CHF 24 each. Add them to the basket?” Before payment, present a complete summary and require an unambiguous confirmation.
When confidence is low, do not bluff. Offer two clearly different interpretations, ask the customer to spell an identifier, show a visual choice or transfer to a person. After two failed attempts on the same detail, change the interaction rather than demanding that the customer repeat themselves louder.
Write for hearing, not reading
Spoken sentences should be shorter than page copy. Put the important distinction early. Avoid presenting more than a few options at once. Do not read navigation, legal terms or product descriptions as one uninterrupted block.
Numbers need special care. A price, date, dosage, dimension or address should be repeated in an easy-to-verify format and displayed where possible. A transcript helps customers review what the system understood and supports people who cannot hear the response.
Use multimodal confirmation for commercial decisions
A screen can show product image, variant, quantity, total price, delivery address and cancellation path simultaneously. Voice must present them sequentially, relying on memory. Combining the two modes is often safer and faster.
Let voice initiate and narrow; let the screen confirm and compare. A phone can display a basket after a spoken reorder. A smart display can show appointment times. An email or message can provide a review link without completing the transaction automatically.
Never use the presence of a screen as permission to hide information in small print. The visual confirmation should make the consequential details prominent and operable by keyboard, touch and speech.
Plan authentication before accepting commands
A voice heard by a shared device is not necessarily the account holder. Homes, workplaces and vehicles contain several speakers, and recorded or synthetic voices add further risk. Treat a spoken command as input, not proof of identity.
Low-risk tasks may require a signed-in device. Viewing sensitive status, changing an address or making a purchase may require a second factor, a voice PIN supplied by the platform or confirmation in an authenticated app. Amazon’s current Shopping Actions documentation describes checks including valid payment method, delivery address and voice-PIN authorisation.
Risk controls should scale with consequence. Do not speak private order or health information where other people may hear it. Allow customers to disable voice purchasing or restrict it to adding items to a reviewable list.
Keep platform dependency visible
Assistant ecosystems determine invocation, discovery, certification, identity, payments and analytics. Their policies and supported features can change. Google’s closure of Conversational Actions shows why the business case must survive a platform withdrawal.
Prefer capabilities that strengthen owned channels: accessible semantic HTML, mobile speech input, a well-designed conversational service inside your own app or website, and APIs that can serve several interfaces. Keep product data, customer permission, transaction records and fulfilment logic outside a single assistant where possible.
| Layer | Business should control | Platform may provide |
|---|---|---|
| Customer promise | Eligible tasks, service levels and support | Invocation and interface conventions |
| Product information | Canonical items, prices, availability and terms | Presentation and entity matching |
| Transaction | Order accuracy, fulfilment, refunds and records | Authentication or approved payment flow |
| Consent and privacy | Purpose, retention, customer rights and vendors | Device permission prompts |
| Continuity | Fallback channel, data export and shutdown plan | Platform status and migration tools |
Treat spoken data as personal data
A voice interaction may contain identity, household context, preferences, account information and sensitive details. Transcripts and audio recordings can create a larger data footprint than the underlying order. Decide whether audio needs to be stored at all. A structured intent record may be enough.
For Swiss businesses, the Federal Act on Data Protection applies to AI-supported processing. The FDPIC says users communicating with intelligent language systems have a right to know when they are speaking with a machine and whether their input is used to improve a system or for other purposes. Purpose, functionality and data sources should be transparent.
Map every processor: device platform, speech-to-text provider, language model, analytics, customer service and commerce system. Define retention, access and deletion. Do not reuse transaction speech for marketing or model training without an appropriate basis and clear information.
Build for languages, accents and different speech
Switzerland makes language testing unavoidable. German, Swiss German, French, Italian and English contain regional accents, borrowed product names and different number conventions. Platform language support does not guarantee that your catalogue, people or delivery locations will be recognised.
Test with real customers, including people who speak slowly, rapidly, with speech impairments or in a non-native language. Include background noise and ordinary microphones. Track failures by intent and language without using the data to exclude difficult users.
Always provide equivalent non-voice operation. Some people cannot or do not want to speak. Others cannot hear a response, lack privacy or are in a noisy location. Accessibility means adding choices, not replacing one mandatory interface with another.
Run a narrow pilot with operational metrics
Choose one frequent task with a clear baseline, such as checking status or repeating a standard order. Prototype the conversation before integrating payment. Test recognition, repair and handoff. Observe people rather than asking whether voice sounds innovative.
| Measure | What it reveals | Guardrail |
|---|---|---|
| Task completion | Whether voice helps finish the intended job | Do not count accidental or corrected orders |
| Correction turns | Recognition and dialogue quality | Review by language and task |
| Time to completion | Convenience versus existing interface | Include authentication and review time |
| Escalation rate | Where automation stops being useful | Measure successful human resolution |
| Order error and cancellation | Commercial harm from misunderstanding | Set a strict stop threshold |
| Repeat voluntary use | Whether customers find sustained value | Exclude forced voice-only journeys |
A pre-launch checklist
- The task is easier to speak than tap.
- The object, quantity and account can be resolved reliably.
- Consequential details receive explicit confirmation.
- A visual or text transcript is available where practical.
- Authentication reflects the consequence of the action.
- Two failed repair attempts trigger another route.
- Customers can cancel, correct and reach a person.
- Audio, transcripts, purposes and vendors are documented.
- Real accents, languages, disabilities and environments were tested.
- The business can continue if the platform removes the feature.
The opportunity is narrower—and better—than the hype
Voice is unlikely to become the ideal interface for browsing a complex catalogue or comparing consequential products. Sound is temporary. Commerce often depends on inspection, memory, evidence and precise consent.
The non-commodity opportunity lies in recognising the small moments where hands and eyes are occupied, typing is difficult or repetition is wasteful. A familiar reorder, status check or simple booking can become noticeably easier. That improvement may matter greatly to a specific customer even if it never becomes a major sales channel.
Start with accessibility and task evidence, not a smart-speaker strategy. Keep the transaction visible and reversible. Own the operational logic and fallback. Voice commerce earns its place when the customer finishes a known task with less effort and no loss of understanding.
For a small business, that may mean never launching a public assistant at all. Improving product identifiers, semantic form labels, order-history APIs and customer-service handoffs can make the existing website work better with speech tools customers already use. Those investments remain useful when devices and platforms change. They also improve ordinary search, support and operations. The durable asset is therefore not a branded voice persona. It is clean information and a well-defined transaction that can be expressed through voice, screen or human service without changing its meaning.
A narrow voice use case is easier to defend when it supports a clear strategy. Test it against the business model rather than the hype cycle, define the customer through niche research, and use content-demand evidence to identify the spoken questions customers actually repeat.
Official standards and platform documentation
- W3C WAI: Speech recognition — accessibility benefits and implementation foundations.
- W3C WAI: Keyboard accessible — support for speech and other input technologies.
- Google: Conversational Actions sunset — official closure timeline and affected developer features.
- Amazon Alexa: Shopping Actions — current product, confirmation and authorisation flow.
- FDPIC: AI and data protection — Swiss transparency and control requirements for AI-supported processing.



