hardwareESP32BLE

Building a 12g AI Device: What We Learned Making Todovo

12 min readTakahiro Torii

We're launching Todovo on Kickstarter July 1. This is the honest story of building it — including the parts we got wrong.

Why We Built Hardware

The obvious question: why not just build an app?

We tried. For the first six months, Todovo was a widget on the iPhone lock screen. You could tap it, speak a task, and it would categorize and file it automatically. The AI worked well. The problem was the hardware — the phone itself.

Pulling your phone out takes 7–15 seconds. In that time, your working memory has already started decaying. We ran our own experiment: tracked every moment where we thought "I should capture this" for two weeks. 86 moments where we failed to capture because the phone was inconvenient.

The insight: the bottleneck isn't AI. The bottleneck is the physical act of getting a screen in front of your face. So we decided to build dedicated hardware.

The Constraints We Designed For

Before writing a single line of firmware, we wrote down the constraints:

  • Size and weight: Had to be pocketable without thinking about it. Target: under 15g, smaller than a car key fob.
  • Battery: Had to be measured in months, not days. If you need to charge it, you'll forget to charge it. Target: 3 months typical use.
  • Latency: From button press to task appearing in the app, under 3 seconds in normal conditions.
  • Cost: End-user price under $100. Kickstarter Early Bird under $70.
  • Social invisibility: No screen. No speaker. No glowing lights during capture.

Choosing the Microcontroller

We evaluated four MCU options:

OptionPower draw (deep sleep)BLE 5.0Price (qty 100)Verdict
ESP32-C3~5µAYes$1.80Strong candidate
ESP32-S3~7µAYes (5.0)$2.10Strong candidate
nRF52840~2µAYes (5.0)$3.20Too expensive for BOM
STM32WB~2.2µAYes (5.0)$3.40Too expensive

We went with ESP32-S3 for three reasons:

  1. The built-in neural network acceleration gave us options for on-device preprocessing of audio before transmission — reducing data sent over BLE.
  2. The development ecosystem is large enough that we could move fast. ESP-IDF is well-documented.
  3. The price point worked. At $2.10 in initial quantities, we had enough BOM budget for a decent microphone and rechargeable battery.

What we'd do differently: The nRF52840 would have given us better sleep power numbers and a more mature BLE stack. If we were building for longer battery targets (6+ months), we'd probably switch.

The Microphone Problem

Audio quality is everything in a voice device. We went through four microphone options before finding one that worked.

MEMS vs. electret: We started with an electret microphone. It was cheap and produced excellent audio quality in a lab. It was terrible in real-world conditions — too sensitive to handling noise, too directional, and the analog front end added circuit complexity.

We switched to a MEMS digital microphone (PDM interface). MEMS microphones have better power consumption, smaller size, and less handling noise. The tradeoff is slightly lower frequency response at the low end — not a problem for voice. Key specs: SNR 66dB, sensitivity -26 dBFS, current draw 650µA active / 10µA sleep.

The positioning problem: Where you put the microphone matters as much as which microphone you choose. Our first prototype had the mic on the PCB, recessed in the enclosure. The enclosure created a resonance chamber that added a consistent mid-frequency peak — voices sounded "boxy." We moved the mic to the edge of the PCB with a port in the enclosure and the problem disappeared.

BLE Architecture: Why We Chose BLE Over Wi-Fi

The original prototype used Wi-Fi. Press the button, the device wakes up, connects to the network, sends the audio, disconnects. The problem: Wi-Fi association takes 2–4 seconds before you've sent a single byte. That alone blew our 3-second latency target.

BLE (Bluetooth Low Energy) maintains a persistent connection with the phone when in range. Button press → BLE packet → phone receives in ~20ms. The phone then handles the network call to the AI backend.

What "3 seconds" actually looks like:

T+0ms:    Button press detected (interrupt-driven, not polled)
T+20ms:   BLE notification sent to phone
T+50ms:   Phone app receives audio start signal
T+0-2000ms: Audio streaming over BLE (1-3 seconds of speech)
T+50ms:   Button release -> BLE end-of-audio signal
T+500ms:  Phone sends audio to AI backend (Whisper)
T+1200ms: AI transcription + categorization complete
T+1250ms: Task written to database
T+1300ms: Phone UI updates

Total: 2.2–3.5 seconds depending on speech length and network conditions.

BLE audio streaming: BLE isn't designed for audio. We encode audio at 16kHz, 16-bit using IMA-ADPCM compression to get it under BLE's practical throughput limit. The compression is done on-device (ESP32-S3's processor handles this comfortably), which reduces packet count and improves reliability.

Power Architecture: How We Got 3 Months

Battery life on a small device is mostly about what you do when not active.

The math: A 150mAh LiPo at 3.7V = 555mWh. If we assume 10 captures per day at 3 seconds each = 30 seconds active per day. If the rest of the time (23 hours, 59.5 minutes) draws 5µA @ 3.3V = 16.5µW, that's 396µWh/day for idle. 555mWh / (396µWh + active_energy) ≈ 3 months.

The key is the deep sleep current. ESP32-S3 with our configuration: ULP (Ultra-Low Power coprocessor) monitors the button GPIO while the main CPU is in deep sleep. Wake-on-interrupt from button press. This gets us to ~12µA deep sleep, including quiescent current from the PMIC and microphone sleep mode.

What consumes power unexpectedly:

  • The BLE connection itself, even when idle, keeps the RF powered at intervals. We tuned the connection interval to balance latency vs. power. At 100ms connection intervals, average BLE idle current is ~8µA.
  • The LDO voltage regulator. We replaced a generic AMS1117 (quiescent current ~5mA) with a TPS62840 (quiescent current ~60nA). This alone saved ~2 weeks of battery life.
  • Leakage in the charging circuit. We added a load switch to disconnect the charger circuit when not actively charging.

The Enclosure: 12g Is Harder Than It Sounds

ComponentMass
PCB (35mm × 22mm × 1mm)1.8g
ESP32-S3 module2.1g
Battery (150mAh)3.2g
Button mechanism0.6g
MEMS microphone0.1g
Enclosure (PA12, SLS)4.2g
Total12.0g

Button feel: We landed on a 0.15N actuation force with ~0.3mm travel using a SMD tactile switch (C&K KSC series). Drop tolerance target was 1.5m onto concrete. The PA12 nylon SLS enclosure absorbs impact better than injection-molded ABS.

The AI Stack

We don't do on-device AI. We considered it, but for natural language understanding, contextual categorization, and intent extraction — on-device models at our size/power budget would produce meaningfully worse results than cloud inference.

What happens when you press the button:

  1. Device streams audio over BLE to the iOS app
  2. App sends audio to our cloud backend
  3. Backend runs Whisper for transcription (fastest available model at quality threshold)
  4. Transcribed text goes to GPT-4o with a system prompt that extracts: task description, deadline, project (inferred), priority (inferred), contact (if mentioned)
  5. Structured task object returned to app
  6. App writes to user's task database (synced to OmniFocus, Things, Notion, or Todoist)

Latency breakdown: Whisper transcription is ~300–500ms for 2–3 second audio clips. GPT-4o structured extraction is ~400–600ms. Network round-trips add ~100ms each. Total backend time: ~900ms–1200ms.

What Failed

  • Wi-Fi first (failed): The first three prototypes assumed Wi-Fi. The latency was unusable. Switching to BLE cost us six weeks.
  • Capacitive touch (failed): Our first button design was capacitive. In a pocket, fabric contact triggered captures. We went back to a physical button.
  • Onboard Whisper (failed): We spent three weeks trying to run a quantized Whisper tiny model on-device. The model runs, but transcription quality was unacceptable for non-native English speakers and any proper nouns.
  • Injection molding for prototypes (wrong call): We spent money on injection mold tooling before we'd finalized the design. We changed the design. The mold was wasted. SLS 3D printing for prototyping, injection molding only when the design is locked.

What We'd Tell Other Hardware Founders

  1. Do the BOM math before you design, not after. We almost couldn't hit $99 retail because we hadn't priced components early enough.
  2. Power consumption is 90% about sleep mode. Optimize what the device does when idle.
  3. The first mold is always wrong. Don't pay for injection tooling until you've frozen the mechanical design.
  4. BLE is harder than Wi-Fi to debug but better for battery and latency. Worth it.
  5. On-device AI sounds cool, but cloud inference is usually better. Especially for NLP. The device is dumb; the intelligence is in the backend. Design accordingly.

We're launching July 1 on Kickstarter. If you're building something similar — or if you've spotted errors in our technical approach — we'd genuinely like to hear from you: feedback@todovo.ai

Join the waitlist at www.todovo.ai →

Related: The Hidden Cost of Task Capture | GTD's Capture Step: The Phase Everyone Gets Wrong | Todovo vs PLAUD Note vs Apple Watch