
Private WebML Demo
Classify photos and chat with LLMs running entirely in your browser tab.
Try live inference →
EdgeBites designs, deploys, and operates AI where your data is born — on-device inference, microcontroller integration, and production edge software. No cloud round-trip required.

Classify photos and chat with LLMs running entirely in your browser tab.
Try live inference →
Machine-filtered edge-AI, MCU, and NPU headlines — hourly.
Read the news →
Short notes from the lab — deployments, measurements, things learned the hard way.
Read the blog →
Architecture, deployment, optimization, microcontrollers, and security — edge AI from sensor to insight.
Overview · Deployment · Optimization · Microcontrollers · Security
Prefer a reader? Micro-Blog RSS feed · Full reading setup at rss.edgebites.com · Fediverse: social.edgebites.com
Edge AI systems for places the cloud can't reach — factory floors, farm fields, wearables, vehicles, and remote sites. We take you from sensor to insight.
Start with the BOM, not the model. Hardware selection, latency and power budgets, costed before silicon.
INT8 quantization, pruning, distillation, NPU delegation — every accuracy point accounted for.
OTA with health gates, staged rollouts, automatic rollback. Proven on 1GB hosts — like this one.
Secure boot, encrypted artifacts, mutual TLS. Designed in, not patched on.
A shippable path from notebook to device — the same pipeline discipline we run ourselves.
This site is served from a 1GB Debian host using exactly this discipline. Ask for a deployment review →
Precise optimizations at every level — from sensor hardware to edge terminals to cloud servers. No watt or millisecond wasted.
Spend microwatts, not milliwatts: adaptive sampling, duty cycling, wake-on-event.
Shrink models 4× with INT8, prune dead weights, delegate to NPU/DSP/CMSIS-NN.
Cache, batch, forward: MQTT/CoAP/LoRaWAN tuned per link budget.
Batch inference, pack accelerators, degrade gracefully when nodes drop.
Autoscale to zero, cache artifacts, burn preemptible watts for batch.
Power traces, latency budgets, accuracy deltas — unmeasured is unoptimized.
The efficiency metric for the LLM era: useful tokens per joule, from battery-powered MCUs and sensors to edge terminals, cloud servers, and GPUs. Right-size the model, quantize aggressively, batch where latency allows — and put each token on the cheapest watt that meets the deadline.
Real models running 100% on your device — vision and text inference with nothing uploaded. Pixels and prompts never leave this browser.
Probing…
Description appears here after you load the VLM and pick an image…
Description:
Community results load here for the selected model + context.
Browse Google AI Mode, then push the page here with the bridge icon (or Ctrl+Shift+E). The local WASM filter renders text windows + prompt and removes side panels — everything stays in this window, nothing leaves your browser.
New thread — open Google AI Mode, then push the page here with the bridge icon (or Ctrl+Shift+E).
Blocked: scripts, trackers, nav, side panels, ads — only text windows + live prompts. Nothing leaves your browser.
Fallback path only. Normal flow needs no ID: install edgebites-ai-bridge.zip (unzip → chrome://extensions → Developer mode → Load unpacked), open a Google AI Mode tab, click the bridge icon or press Ctrl+Shift+E.
Push flow: install edgebites-ai-bridge.zip, open Google AI Mode, click the bridge icon (or Ctrl+Shift+E) — the page renders here, panels removed. Direct page fetch is blocked by Google CORS, so the companion extension carries the HTML inside your own browser. This page never opens new windows itself. Extractor: /assets/ai-extract/pkg/ai_extract_bg.wasm.
Milliwatt-scale intelligence — sensing, fusion, and control on microcontrollers.
Vibration, acoustic, and vision anomaly detection that sips power on Cortex-M and ESP32-S3.
IMU, environmental, and GNSS fused on-device; only events use the uplink.
Buffer, store-and-forward, translate: MCU-to-Linux plumbing.
Where it matters most, at the edge — protection and prioritization computed on-device, not in a distant cloud.
Every node carries keys and attests firmware. No trust, not even on LAN.
Mutually authenticated, encrypted transit — our own mail and admin ports run TLS-only.
Alerts jump the queue; telemetry waits. Lossy links stay useful.
Signed artifacts, staged rollout, auto-revert on trouble.
Local anomaly models flag intrusions in milliseconds — offline included.
Rate limits and edge filtering keep fleets answering when the internet won't.
We build in the open — tooling, examples, and building blocks under open licenses. Start here:
github.com/EdgeBites — all repositories →
Reproducible single-host edge stacks — the configs serving this page.
Reference MCU integrations published with measured latency and power.
RSS, federation, and mail glue for a presence you fully own.
Short notes from the lab — deployments, measurements, and things we learned the hard way. Also available as RSS.
Loading posts…
Edge-AI, MCU, and NPU headlines from across the web — machine-filtered for signal, refreshed hourly. Full reading in FreshRSS.
Loading headlines…
Download OPML to import this feed list into any reader.
Tell us about your edge project — what senses, what decides, what constrains it. We reply by email, usually within two business days.
Prefer the fediverse? social.edgebites.com