EdgeAI Innovations Inc.
Moorcheh
Memory Engine for Memanto
Production-grade memory for your agent fleet, managed by Memanto, powered by Moorcheh. Product description and pricing for the Google Cloud Marketplace listing.
Memanto and Moorcheh: How They Work Together
Memanto is the free, open-source (MIT) companion memory agent trusted by thousands of developers. It runs alongside your AI agents and curates, reconciles, consolidates, briefs, and provisions what they remember.
Moorcheh is the production search, retrieval, and intelligence engine underneath. It indexes and recalls every memory Memanto manages - governed, fast, and cost-efficient at fleet scale.
Memanto stays free and open source. Moorcheh is how teams run that same memory architecture in production.
Where This Listing Fits
Moorcheh is available three ways. If you are evaluating, start free. This listing is the scale-up path.
Free local image
Run Moorcheh on your own machine for development and evaluation at no cost.
Moorcheh Cloud
Community and Production tiers at moorcheh.ai for teams getting started and early production.
This listing
Dedicated single-tenant Moorcheh - capacity guarantees, committed SLAs, and isolation at the Google Cloud project level.
Why Moorcheh is Different
Purpose-built for Memanto
Not a general-purpose database with a memory API attached. Designed alongside Memanto and published together in peer-reviewed literature - deterministic retrieval under 90 ms from a single query, with no ingestion delay (arXiv:2604.22085).
Replaces your vector database and reranker
No separate vector store to license, no embedding index to maintain, no external reranking fee. MAIR benchmark: quality comparable to full-precision systems at substantially lower latency (arXiv:2601.11557).
No in-memory index, no RAM bill
Exhaustive search over compact binary representations - 32× smaller than float32 - removes the always-on ANN index and the RAM it consumes.
No idling servers
Because there is no index to keep warm, Moorcheh scales genuinely to zero. Capacity is consumed when queries run.
Your cloud discount stays yours
Provisioned against your Google Cloud billing account. Google invoices you directly at your negotiated rates. Our licence fee is our entire fee.
A project of your own
Dedicated single-tenant Google Cloud project. You choose the region at onboarding so residency is settled before anything is provisioned.
Operated for you - 500+ assets
We install, configure, patch, upgrade, monitor, and support every coordinated Google Cloud asset. Installation included in every plan.
Deterministic under load
Same correct answer every time; accuracy does not degrade as concurrency rises. Memanto on Moorcheh: 89.8% LongMemEval, 87.1% LoCoMo (arXiv:2604.22085).
Deterministic by default
Remembering, recalling, migrating, and scoping never touch a language model. Only genuine reasoning - answering, conflict resolution, briefing, consolidation - invokes one.
Governed by default
Per-agent scoping, audit logs, and contradiction reporting on every plan. No SSO tax, no enterprise upsell for compliance features.
The Cost Stack You Remove
A production memory deployment is rarely one line item. Moorcheh collapses that stack into one engine. When you compare this listing against a memory platform subscription alone, you are comparing against one line of a five-line bill.
| Typical production memory stack | With Moorcheh |
|---|---|
| Memory platform subscription | Included |
| Vector database - licence, storage, read/write units | Not required |
| Reranking service - per-query fee | Not required |
| Always-on RAM for the in-memory ANN index | Not required - scales to zero |
| Index build and re-index compute on ingest | Not required - no indexing |
| Model tokens on every memory read and write | Agentic calls only |
| Vendor margin on your infrastructure | None - Google bills you directly |
Dedicated. Operated. Billed to you.
Every plan provisions a dedicated, single-tenant Moorcheh deployment in a Google Cloud project of its own, created and operated by our team. Your deployment is billed to your own Cloud Billing account, so Google invoices you directly for infrastructure at your negotiated rates and your committed spend applies. Plans cover your Moorcheh licence, installation, capacity entitlement, updates, and support. We take no margin on your infrastructure.
Choose a plan based on throughput, memory estate capacity, and resilience posture. Every plan runs the identical engine. Upgrading is a configuration change, not a migration: no re-indexing, no re-ingest, and no downtime for your agent fleet.
You are alerted at 80% of your compressed estate envelope; writes stop at 100%. Reads and recall continue. A full export is available on request at any time. On cancel, we export first, confirm you have it, then tear the deployment down.
Bring Your Own Model
Every plan uses your own Vertex AI endpoint or API keys. Memory never passes through models we control, there are no hidden inference charges, and token costs stay on your own cloud account. Because Moorcheh delegates as much work as possible to deterministic operations, memory traffic can grow without your token bill growing alongside it.
How You Get Started
Not self-serve. Your deployment is more than 500 coordinated Google Cloud assets and we build it for you.
- 01
You subscribe
Choose a plan. Nothing is charged yet.
- 02
Onboarding call
Within 1-2 business days: billing account, region, residency, and sizing.
- 03
We build and validate
Typically 3-5 business days. Dedicated project, full estate, end-to-end test.
- 04
Handover - trial starts
14-day trial begins when the deployment is live, not when you subscribed.
- 05
You decide
Continue and licence charges start day 15 - or cancel, export, and tear down with no licence fee.
The trial waives our licence fee, not your Google Cloud costs. Google charges infrastructure during the trial whether or not you continue. On Launch this is modest (scale-to-zero when idle). On Scale and Performance it is higher. We give you an estimate at the onboarding call, before anything is provisioned.
Launch · Scale · Performance
All prices USD. Google Cloud infrastructure is billed by Google directly to your Cloud Billing account, at your rates and against your committed spend.
Launch
or $3,990 / yr (save $798)
- Memanto agents25
- API calls / mo1.5M
- Agentic calls / mo20K
- Max agents250
- Estate capacity25 GB compressed
- SLA99.5%
- Warm capacityScale-to-zero; brief cold starts possible
- ThroughputUp to ~50 QPS burst
- SupportEmail & Discord
Overage
Extra agent$3.00 / agent / mo
API calls$20 per 1M
Agentic calls$6 per 1K
Scale
or $9,990 / yr (save $1,998)
- Memanto agents75
- API calls / mo8M
- Agentic calls / mo75K
- Max agents1,000
- Estate capacity100 GB compressed
- SLA99.9%
- Warm capacityWarm instance pool during business hours
- ThroughputUp to ~250 QPS
- SupportPriority response
Overage
Extra agent$2.50 / agent / mo
API calls$15 per 1M
Agentic calls$4 per 1K
Performance
or $29,990 / yr (save $5,998)
- Memanto agents200
- API calls / mo30M
- Agentic calls / mo250K
- Max agents5,000
- Estate capacity500 GB compressed
- SLA99.95%
- Warm capacityProvisioned minimum instances 24/7, no cold starts
- Throughput~1,000+ QPS with priority concurrency
- SupportNamed contact + private Slack
Overage
Extra agent$2.00 / agent / mo
API calls$10 per 1M
Agentic calls$3 per 1K
Included on every plan
- Dedicated single-tenant deployment
- Billed to your own Cloud Billing account
- Installation and setup included
- Deterministic information-theoretic retrieval
- Serverless scale-to-zero architecture
- Bring your own model (Vertex AI or API key)
- Per-agent scoping within your estate
- Audit logs & contradiction reporting
- SSO / SAML
- Automated updates
- 14-day trial (from handover)
- Marketplace billing (draws down committed spend)
Annual Subscriptions
Ten months' price for twelve months of service. Entitlements, overage, capacity, and SLA match the monthly plan. Metered overage still bills monthly in arrears.
| Plan | Monthly | Annual | You save |
|---|---|---|---|
| Launch | $399 / mo | $3,990 / yr | $798 |
| Scale | $999 / mo | $9,990 / yr | $1,998 |
| Performance | $2,999 / mo | $29,990 / yr | $5,998 |
What We Measure
The distinction between API calls and agentic calls is why your token bill does not scale with your memory traffic.
Memanto agents
Creating a Memanto agent provisions a dedicated namespace in your deployment. Count reflects provisioned namespaces (peak in the billing period), not how busy the fleet has been. Decommissioning deletes the namespace and stops the count.
API calls
Deterministic operations: remembering, recalling, scoping, listing, migration, and admin. These run entirely inside Moorcheh - no language model, no tokens on your account. Most production fleet traffic falls here.
Agentic calls
Operations that invoke Memanto's agentic functionality: answer, conflict resolution, briefing, consolidation, and other reasoning. You are billed by EdgeAI for the agentic call and by your model provider for tokens. No inference margin in our fee.
Need Something Larger?
Enterprise is available as a private offer for unbounded agent counts and estate size, custom resilience and recovery targets, a bespoke SLA, negotiated commercial terms, and named engineering support.
EdgeAI Innovations Inc. · support@moorcheh.ai · moorcheh.ai