If you're building or sourcing an AI toy, one decision shapes almost everything else: where does the AI actually run? It can run in the cloud, inside the toy, or a bit of both.
There's no answer that fits every product. There is one that fits yours, once you know your buyers, your markets and your margins. This guide lays out the trade-offs in plain terms, so you can decide before you pay for samples.
It's written for three kinds of readers: retailers and distributors choosing what to stock, brands and IP owners turning a character into an AI toy, and toy companies adding AI to an existing line.
For most AI toys, hybrid is the safest choice: keep the basics on the toy and send only the hard conversations to the cloud. Wake-up, simple commands, sound effects and stored stories need to work instantly and everywhere. Open-ended chat needs a large model, and large models live on servers.
There are exceptions. A simple bedtime storyteller can go fully offline. A premium companion that stays connected at home can reasonably go cloud-first.
| Architecture | Best for | Main upside | Main catch |
|---|---|---|---|
| Offline | Simple storytellers, toddler toys, travel toys | Works anywhere, and voice processing stays on the toy | Limited open conversation and less room to grow |
| Cloud | Premium companions, language learning, knowledge Q&A | Richest conversation, easy to improve after launch | Needs a connection, has a running cost, and moves more data off the toy |
| Hybrid | Most branded and retail products | Reliable basics, plus deeper chat when online | More engineering: two paths to design and test |
The rest of this guide shows how to get from that summary to a decision for your own product.
The difference comes down to where speech and language are processed: on the toy's own chip, on a remote server, or split between the two.
In a cloud setup, the toy is mostly a microphone, a speaker and a connection. The thinking happens elsewhere.
What you gain: the richest conversation, and the ability to improve the toy after launch without touching the hardware.
What you give up: it needs a connection, every conversation has a cost, and the child's voice leaves the toy.
In an offline setup, speech recognition and responses run on the toy's own chip. No connection is needed once it's set up.
What you gain: it works in the car, on a plane and at grandma's house. Responses are fast, and less data has to leave the toy.
What you give up: the chip, memory and battery set a hard ceiling. Offline toys usually handle wake words, commands and prepared content well, but free-flowing conversation is limited. Updates mostly arrive as firmware.
A hybrid toy splits the work by difficulty. The toy handles wake-up, commands, sound effects, stored stories and a fallback mode. The cloud handles open-ended chat, knowledge questions and fresh content.
The part that matters most is the fallback. When Wi-Fi drops, a hybrid toy should still behave like a toy, not go silent.
Here's how the three compare on the factors buyers and engineers argue about most.
| Factor | Offline | Cloud | Hybrid |
|---|---|---|---|
| Internet needed | No | Yes | Only for advanced features |
| Conversation depth | Limited and structured | Open and flexible | Open when online, basic when offline |
| Voice data leaving the toy | Minimal by design | Yes, by design | Depends on what you route to the cloud, and you can limit it |
| Running cost to you | Low | Recurring usage cost | Lower than cloud-only, since routine requests stay local |
| Response speed | Fast | Depends on the network | Fast for basics |
| When Wi-Fi drops | Unaffected | Toy goes quiet or limited | Falls back to on-toy features |
| Improving after launch | Firmware updates | Server-side, easy | Both |
| Engineering effort | Fitting AI into small hardware | Building and running a backend | Highest: two paths to design and test |
Hybrid isn't free. You're designing, testing and maintaining two paths instead of one. It's worth it when the product needs both reliability and depth, which most branded AI toys do.
When the user is a child, where the voice data goes is a product-defining choice, not a footnote.
Privacy rules for children differ by market, but the principle is the same everywhere. In the US, the FTC's COPPA Rule puts parents in control of the personal information collected from children under 13. In the EU, the GDPR says children merit specific protection for their personal data, in particular when data is collected through services offered directly to them.
Regulators are also looking closely at AI that talks like a friend. In September 2025 the FTC issued information orders to seven companies that run companion-style chatbots, asking how they test for harm to children and teens and how they use or share personal information from conversations. Those orders went to chatbot platforms, not toy makers, but they show where attention is heading.
Independent testers have been looking at toys directly. In its 2025 Trouble in Toyland report, U.S. PIRG Education Fund tested four AI chatbot toys. It reported that some had limited or no parental controls, and that these toys can record a child's voice and collect other sensitive data. Separately, a University of Cambridge study of young children and AI toys found that many generative AI toys have privacy practices that are unclear or missing important details.
One caution: processing on the toy means less data has to leave it, but that's a design choice, not a guarantee. What the toy stores, and what any companion app collects, still matters.
What this means for your design:
Cloud AI turns part of your product into a service, so every conversation costs something for as long as the toy is in use.
Hardware is paid once. Cloud usage keeps going, and it grows with how much children talk to the toy and how long they keep it. There are three common ways to cover that cost.
| Approach | Who pays | Watch out for |
|---|---|---|
| Build it into the retail price | The brand | Margins shrink if real usage beats your forecast |
| Charge a subscription | The consumer | Buyers can feel misled if it isn't clear before purchase |
| Free basics, paid extras | Shared | You need a clear, honest line between free and paid |
Hybrid helps here because routine interactions never touch the server. Only the requests that need a large model carry a cost.
Do the math before you design. A simple model is enough:
Ask any supplier for the real cloud cost per active toy under your usage assumptions. That number should be in the quote, not a surprise six months later.
A cloud-only toy goes quiet when the connection does, and children notice immediately.
Think about where toys actually get used: car rides, flights, grandparents' homes, classrooms with restricted networks, and the first ten minutes out of the box before anyone has paired anything. Each of those is a place where a cloud-only toy can fail.
The fix is to decide in advance what must work offline and what can wait for a connection.
| Must work offline | Fine to need a connection |
|---|---|
| Wake-up and basic voice commands | Open-ended conversation |
| A core set of songs and stories | Knowledge and homework questions |
| Sound effects and simple reactions | New stories and content packs |
| Volume, sleep and parental controls | Deeper personalization |
A toy that's fun straight out of the box, before pairing, also removes the biggest day-one friction for parents.
Match the architecture to the experience: simple and portable favors offline, deep and evolving favors cloud, and most products in between favor hybrid.
| Product type | Usually best | Why |
|---|---|---|
| Early-childhood storyteller or sleep plush | Offline-first | Content is mostly fixed, bedtime has to be reliable, and little data is needed |
| Character IP companion | Hybrid | Voice and basic reactions must always work, with deeper chat when online |
| Educational or multilingual toy | Hybrid or cloud | Broad knowledge and many languages help, but core lessons should stay local |
| Robot or motion toy | Hybrid | Movement reactions need low latency, conversation can use the cloud |
| Retail-ready, off-the-shelf product | Hybrid, local-first | It works out of the box, so there's less setup friction for buyers |
| Premium, always-connected companion | Cloud with local fallback | Best conversation quality, and buyers expect an ongoing service |
A ready-made product is the low-risk way to test a market. A custom build makes sense when you need your own character, voice, knowledge base or architecture, and want those to be hard for competitors to copy.
Work through six questions in order, and the right architecture usually becomes obvious by step four.
Not sure where your product lands? Tell us about it in the form below and we'll suggest an architecture.
Architecture is a software decision, but it lands on real hardware, and our three AI modules cover the main product shapes. Full disclosure: we make these modules. Here's what each one does, using our own spec sheet.
| Module | What it adds | Core hardware | Typical product |
|---|---|---|---|
| AI Audio Module (ZN-AI-001) | Voice chat, continuous dialogue, interruption, voice wake-up | Dual microphones, 800mAh battery, 3W speaker, Type-C charging | Plush toys, dolls, storytellers |
| AI Audio Display Module (ZN-AI-002) | Everything above, plus on-screen expressions and emotion recognition | 0.71-inch screen, 10 built-in expressions | Expressive robots, mascots |
| AI Audio Motion Module (ZN-AI-003) | Voice plus movement, touch sensing and intent recognition | Drives up to 3 motors, touch-sensing board, Bluetooth audio | Robot pets, action figures |
Each comes in Basic, Perception and Vision versions, with Wi-Fi, BLE, BT and 4G connectivity.
The architecture is yours to choose. Offline, hybrid and cloud setups, subscription-free models and over-the-air updates can all be built to project requirements.
The modules connect to multiple large models. Our spec sheet lists DeepSeek-V3, ERNIE Bot, Qwen, iFlytek Spark and others. Character personality, voice and knowledge can be customized around your IP.
For details on any of this, use the form below and tell us your product and target market.
We run one process from first idea to shipment, backed by in-house R&D and a manufacturing base of more than 40,000 square meters. The steps, as listed on our AI toy solution page:
A good supplier answers these eight questions clearly and in writing, and a vague answer to any of them is a warning sign.
We're happy to answer all eight for our own modules. Just ask in the form below.
Not always. A cloud-based toy needs a connection for AI conversation. An offline toy runs its main functions on the device. A hybrid toy keeps the basics working without Wi-Fi and uses the cloud for advanced features.
For open-ended conversation, usually not, because the toy's chip, memory and battery limit what can run on it. For wake words, commands, stored stories and simple reactions, offline is plenty.
Yes. That's a hybrid design. Local processing handles the basics and fallback, and the cloud handles complex conversation and new content.
No. Less data has to leave the toy, but what the toy stores and what any companion app collects still matter. Ask your supplier exactly what is collected, where it goes and whether it can be deleted.
Not necessarily. Subscriptions usually exist to cover cloud running costs. Designs that keep routine features on the toy, or price cloud use into the product, can avoid them.
In many cases, yes. It depends on the space available for the microphone, speaker, battery and module. Send us your design and we can review it with you.
The best architecture is the one that makes your toy reliable on day one, affordable to run for years, and easy to explain to a parent.
Not sure which setup fits your product?
Tell us about it in the form below. The more you share, the better our recommendation: toy type, target age, markets, languages, features that must work offline, expected volume and launch date.