Moorcheh Edge on Arduino UNO Q vs Ventuno Q: 1536-D Vector Benchmark
We compared two ARM64 edge boards running the same Moorcheh Edge workload: Arduino UNO Q (about 3.6 GB RAM) and Ventuno Q (about 15 GB RAM). Each board ingested and searched 1536-dimensional vectors in vector mode only — no text pipeline, no embedding model, no LLM. Queries were raw float vectors sent to POST /search.
For this benchmark we raised the per-store item limit to 1,000,000 so we could stress-test ingest and search at scale. Typical Moorcheh Edge deployments use a smaller cap suited to kiosk-sized catalogs. Even with Moorcheh's compact storage, 1M vectors is an extreme case; it shows how the two boards behave when the catalog grows.
For the earlier UNO Q RAG pipeline breakdown (embed + LLM vs retrieval), see Edge RAG on 4 GB Hardware.
Headline results: At 10k vectors, median search latency is a tie (about 25 ms on both). At 1M, Ventuno Q is about 3× faster on search (269 ms vs 833 ms median) and about 4× faster on full ingest (8 hours vs 32 hours wall clock).
Detailed tables (upload progress, 50 search samples per run, RAM during ingest, and the 50-query search + RAM study at 1M) are in the Google Sheets linked under Full results.
Why Moorcheh Edge makes this benchmark possible
A conventional 1M × 1536 float32 vector store is roughly 6 GB of embedding data alone (1,000,000 × 1536 × 4 bytes), before index structures and runtime overhead. That does not fit the practical memory budget on Arduino UNO Q and is a poor use of Ventuno Q if the goal is edge deployment rather than raw storage.
Moorcheh Edge binarizes vectors at ingest with MIB (Moorcheh Information Binarization) and retrieves with EDM (Enhanced Distance Metric) over those compact forms. In our runs:
- On disk: about 3.2 MiB (10k), about 32 MiB (100k), about 319 MiB (1M)
- In Docker at 1M: about 1.2–1.5 GiB while serving search
The same item counts on both boards confirm that store size is driven by Moorcheh's encoding, not by which CPU runs the server. That small footprint is why we could run a 1M × 1536 stress test on real edge hardware instead of only on a cloud instance.
How we ran the benchmark
We used one Moorcheh Edge server deployment on each board (ARM64, vector-only stack). The client ran on the board against localhost, so timings include HTTP JSON plus server work.
Procedure for each catalog size (10k, 100k, 1M):
- Clear the store (fresh run at that scale).
- Upload random unit vectors in batches until the target count is reached; record wall time and throughput.
- Sample host and container memory during ingest.
- Run 5 warmup searches, then 50 timed searches with
top_k = 5, fixedseed = 42, and report min, median, mean, p95, and p99.
After the 1M run completed on each board, we ran a separate 50-query study on the loaded 1M catalog: before and after each search, record latency and RAM (host used memory and Docker stats for the Moorcheh container).
Both boards received the same software and the same test parameters; only hardware resources differ.
Full results
Every run is documented in two public Google Sheets — each contains the summary, per-query search latencies, upload progress, RAM samples during ingest, and the 50-query RAM worksheet:
Test configuration
| Parameter | Value |
|---|---|
| Server | Moorcheh Edge (same deployment on both boards) |
| Store limit (this benchmark) | 1,000,000 items (raised for stress testing) |
| Store mode | Vector (precomputed floats from client) |
| Dimension | 1536 (locked on first upload) |
| Query | Random unit vector, seed 42 |
| Search | top_k = 5, 5 warmup + 50 timed runs |
| Catalog sizes tested | 10,000 / 100,000 / 1,000,000 (separate runs) |
| Device | RAM context (from measurements) |
|---|---|
| Arduino UNO Q | Docker limit about 3.58 GiB; strong pressure at 1M |
| Ventuno Q | Docker limit about 14.93 GiB; same Moorcheh server, more headroom |
Results: store size
On-disk store size was identical on both boards at every scale.
On-disk store size: Moorcheh MIB store vs raw float32 (MiB)
Green: Moorcheh store (identical on both boards). Red: raw float32 embeddings at 1M × 1536 — ~19× larger than the Moorcheh store, before index structures or runtime overhead.
Moorcheh Edge loads the binarized catalog into memory for search; it does not keep full float32 matrices for every vector.
Results: upload (ingest)
| Vectors | Arduino UNO Q | Ventuno Q | Ventuno advantage |
|---|---|---|---|
| 10k | 59.3 s | 29.5 s | about 2.0× faster |
| 100k | about 19 min | about 10 min | about 1.9× faster |
| 1M | about 32.1 h | about 8.0 h | about 4.0× faster |
At 10k, UNO Q ingest stays practical (about 1 minute). At 1M, ingest dominates wall clock: UNO Q needs on the order of 32 hours for a full load; Ventuno about 8 hours.
Peak host memory during ingest:
| Scale | UNO Q peak (MB) | Ventuno peak (MB) |
|---|---|---|
| 10k | 817 | 1,028 |
| 100k | 875 | 1,155 |
| 1M | 2,673 | 2,708 |
Both boards reach about 2.7 GB host use at 1M. UNO Q is near its ceiling; Ventuno still has room below its Docker limit.
Results: search latency (50 runs per scale)
Median search latency by catalog size (ms, top_k = 5, 50 runs)
Tie at 10k; at 1M Ventuno Q is ~3.1× faster (269 vs 833 ms). p99 and per-run samples are in the linked spreadsheets.
Search cost grows with N. At kiosk scale (10k), both boards deliver about 25 ms median, which is suitable for interactive edge retrieval. At 1M, UNO Q is about 0.83 s per query vs Ventuno about 0.27 s. Latency scales roughly linearly with catalog size, consistent with scanning the full in-memory binarized index.
Results: 1M catalog, 50 queries with RAM per search
With 1M vectors already loaded, we measured each of 50 searches together with RAM immediately before and after the request. Typical host RAM stayed around 2.26 GB on UNO Q and 2.11 GB on Ventuno Q.
50-query study on the loaded 1M catalog
Median latency (ms)
Docker memory (% of container limit)
Similar index footprint (~1.2–1.5 GiB) on both boards — Ventuno's extra RAM is headroom, not a larger index. All 50 rows are in the linked spreadsheets.
Per-query RAM barely changes during search because the catalog is already resident. The gap is CPU and system throughput, not index size.
Choosing a board
Arduino UNO Q fits Moorcheh Edge at 10k (and similar) catalog sizes: about 25 ms search, about 1 min ingest, about 3 MiB store. 100k is usable if about 100 ms retrieval is acceptable. 1M is viable as a stress experiment on this hardware but not as a default production target without long ingest times and about 800 ms-class search.
Ventuno Q runs the same Moorcheh Edge stack but is the better match for 100k–1M catalogs: lower search latency at scale and much shorter 1M ingest.
Methodology and limitations
- Search timings include localhost HTTP JSON, not internal Rust-only timers.
- Each run uses the same seed and query vector for repeatability, not a mix of query types.
- We report vector ingest, search speed, and memory only (no text RAG quality or LLM tests).
- Ventuno Q is our name for a higher-RAM ARM64 edge peer in this study.
- Product store limits for typical edge use remain below the 1M cap used here; we raised the limit only to stress-test Moorcheh Edge at extreme scale.
Try Moorcheh Edge
pip install moorcheh-edge
moorcheh-edge up -y
moorcheh-edge status- On-Edge introduction
- Edge product overview
- Research: From HNSW to Information-Theoretic Binarization
Benchmarks run 29 September – 1 October 2026.
Build this architecture today.
Get your API key and start building agentic memory in under 5 minutes.
Get API Key