EdgeAI Innovations Inc.

Google Cloud Marketplace

Moorcheh

Memory Engine for Memanto

Production-grade memory for your agent fleet, managed by Memanto, powered by Moorcheh. Product description and pricing for the Google Cloud Marketplace listing.

Overview

Memanto and Moorcheh: How They Work Together

Memanto is the free, open-source (MIT) companion memory agent trusted by thousands of developers. It runs alongside your AI agents and curates, reconciles, consolidates, briefs, and provisions what they remember.

Moorcheh is the production search, retrieval, and intelligence engine underneath. It indexes and recalls every memory Memanto manages - governed, fast, and cost-efficient at fleet scale.

Memanto stays free and open source. Moorcheh is how teams run that same memory architecture in production.

Differentiation

Why Moorcheh is Different

Purpose-built for Memanto

Not a general-purpose database with a memory API attached. Designed alongside Memanto and published together in peer-reviewed literature - deterministic retrieval under 90 ms from a single query, with no ingestion delay (arXiv:2604.22085).

Replaces your vector database and reranker

No separate vector store to license, no embedding index to maintain, no external reranking fee. MAIR benchmark: quality comparable to full-precision systems at substantially lower latency (arXiv:2601.11557).

No in-memory index, no RAM bill

Exhaustive search over compact binary representations - 32× smaller than float32 - removes the always-on ANN index and the RAM it consumes.

No idling servers

Because there is no index to keep warm, Moorcheh scales genuinely to zero. Capacity is consumed when queries run.

Your cloud discount stays yours

Provisioned against your Google Cloud billing account. Google invoices you directly at your negotiated rates. Our licence fee is our entire fee.

A project of your own

Dedicated single-tenant Google Cloud project. You choose the region at onboarding so residency is settled before anything is provisioned.

Operated for you - 500+ assets

We install, configure, patch, upgrade, monitor, and support every coordinated Google Cloud asset. Installation included in every plan.

Deterministic under load

Same correct answer every time; accuracy does not degrade as concurrency rises. Memanto on Moorcheh: 89.8% LongMemEval, 87.1% LoCoMo (arXiv:2604.22085).

Deterministic by default

Remembering, recalling, migrating, and scoping never touch a language model. Only genuine reasoning - answering, conflict resolution, briefing, consolidation - invokes one.

Governed by default

Per-agent scoping, audit logs, and contradiction reporting on every plan. No SSO tax, no enterprise upsell for compliance features.

Cost stack

The Cost Stack You Remove

A production memory deployment is rarely one line item. Moorcheh collapses that stack into one engine. When you compare this listing against a memory platform subscription alone, you are comparing against one line of a five-line bill.

Typical production memory stackWith Moorcheh
Memory platform subscriptionIncluded
Vector database - licence, storage, read/write unitsNot required
Reranking service - per-query feeNot required
Always-on RAM for the in-memory ANN indexNot required - scales to zero
Index build and re-index compute on ingestNot required - no indexing
Model tokens on every memory read and writeAgentic calls only
Vendor margin on your infrastructureNone - Google bills you directly
Deployment

Dedicated. Operated. Billed to you.

Every plan provisions a dedicated, single-tenant Moorcheh deployment in a Google Cloud project of its own, created and operated by our team. Your deployment is billed to your own Cloud Billing account, so Google invoices you directly for infrastructure at your negotiated rates and your committed spend applies. Plans cover your Moorcheh licence, installation, capacity entitlement, updates, and support. We take no margin on your infrastructure.

Choose a plan based on throughput, memory estate capacity, and resilience posture. Every plan runs the identical engine. Upgrading is a configuration change, not a migration: no re-indexing, no re-ingest, and no downtime for your agent fleet.

You are alerted at 80% of your compressed estate envelope; writes stop at 100%. Reads and recall continue. A full export is available on request at any time. On cancel, we export first, confirm you have it, then tear the deployment down.

Bring Your Own Model

Every plan uses your own Vertex AI endpoint or API keys. Memory never passes through models we control, there are no hidden inference charges, and token costs stay on your own cloud account. Because Moorcheh delegates as much work as possible to deterministic operations, memory traffic can grow without your token bill growing alongside it.

Onboarding

How You Get Started

Not self-serve. Your deployment is more than 500 coordinated Google Cloud assets and we build it for you.

  1. 01

    You subscribe

    Choose a plan. Nothing is charged yet.

  2. 02

    Onboarding call

    Within 1-2 business days: billing account, region, residency, and sizing.

  3. 03

    We build and validate

    Typically 3-5 business days. Dedicated project, full estate, end-to-end test.

  4. 04

    Handover - trial starts

    14-day trial begins when the deployment is live, not when you subscribed.

  5. 05

    You decide

    Continue and licence charges start day 15 - or cancel, export, and tear down with no licence fee.

The trial waives our licence fee, not your Google Cloud costs. Google charges infrastructure during the trial whether or not you continue. On Launch this is modest (scale-to-zero when idle). On Scale and Performance it is higher. We give you an estimate at the onboarding call, before anything is provisioned.

Pricing

Launch · Scale · Performance

All prices USD. Google Cloud infrastructure is billed by Google directly to your Cloud Billing account, at your rates and against your committed spend.

Launch

$399 / mo

or $3,990 / yr (save $798)

  • Memanto agents25
  • API calls / mo1.5M
  • Agentic calls / mo20K
  • Max agents250
  • Estate capacity25 GB compressed
  • SLA99.5%
  • Warm capacityScale-to-zero; brief cold starts possible
  • ThroughputUp to ~50 QPS burst
  • SupportEmail & Discord

Overage

Extra agent$3.00 / agent / mo

API calls$20 per 1M

Agentic calls$6 per 1K

Talk to us
Popular

Scale

$999 / mo

or $9,990 / yr (save $1,998)

  • Memanto agents75
  • API calls / mo8M
  • Agentic calls / mo75K
  • Max agents1,000
  • Estate capacity100 GB compressed
  • SLA99.9%
  • Warm capacityWarm instance pool during business hours
  • ThroughputUp to ~250 QPS
  • SupportPriority response

Overage

Extra agent$2.50 / agent / mo

API calls$15 per 1M

Agentic calls$4 per 1K

Talk to us

Performance

$2,999 / mo

or $29,990 / yr (save $5,998)

  • Memanto agents200
  • API calls / mo30M
  • Agentic calls / mo250K
  • Max agents5,000
  • Estate capacity500 GB compressed
  • SLA99.95%
  • Warm capacityProvisioned minimum instances 24/7, no cold starts
  • Throughput~1,000+ QPS with priority concurrency
  • SupportNamed contact + private Slack

Overage

Extra agent$2.00 / agent / mo

API calls$10 per 1M

Agentic calls$3 per 1K

Talk to us

Included on every plan

  • Dedicated single-tenant deployment
  • Billed to your own Cloud Billing account
  • Installation and setup included
  • Deterministic information-theoretic retrieval
  • Serverless scale-to-zero architecture
  • Bring your own model (Vertex AI or API key)
  • Per-agent scoping within your estate
  • Audit logs & contradiction reporting
  • SSO / SAML
  • Automated updates
  • 14-day trial (from handover)
  • Marketplace billing (draws down committed spend)

Annual Subscriptions

Ten months' price for twelve months of service. Entitlements, overage, capacity, and SLA match the monthly plan. Metered overage still bills monthly in arrears.

PlanMonthlyAnnualYou save
Launch$399 / mo$3,990 / yr$798
Scale$999 / mo$9,990 / yr$1,998
Performance$2,999 / mo$29,990 / yr$5,998

What We Measure

The distinction between API calls and agentic calls is why your token bill does not scale with your memory traffic.

Memanto agents

Creating a Memanto agent provisions a dedicated namespace in your deployment. Count reflects provisioned namespaces (peak in the billing period), not how busy the fleet has been. Decommissioning deletes the namespace and stops the count.

API calls

Deterministic operations: remembering, recalling, scoping, listing, migration, and admin. These run entirely inside Moorcheh - no language model, no tokens on your account. Most production fleet traffic falls here.

Agentic calls

Operations that invoke Memanto's agentic functionality: answer, conflict resolution, briefing, consolidation, and other reasoning. You are billed by EdgeAI for the agentic call and by your model provider for tokens. No inference margin in our fee.

Need Something Larger?

Enterprise is available as a private offer for unbounded agent counts and estate size, custom resilience and recovery targets, a bespoke SLA, negotiated commercial terms, and named engineering support.

EdgeAI Innovations Inc. · support@moorcheh.ai · moorcheh.ai