The 110% Documentation
Welcome to the ultimate deep-dive into the Stateful AI Intelligence architecture driving HuMem.cloud. This document details how we simulate human memory digitally. Our primary philosophy is simple: Humem = Human Memory. The only difference between humem memory and human memory is that one is stored organically, while the other is stored digitally at the edge.
Sovereign State Management & Edge-Native Substrate
HuMem.cloud eliminates centralized database dependencies to achieve sub-millisecond local state access and absolute per-tenant isolation. It embeds retrieval and storage pipelines directly at the edge using Cloudflare Workers AI.
Edge-Native Embedding Engine
The system utilizes the open-source BGE-M3 multilingual model (@cf/baai/bge-m3) running serverless on Cloudflare's global GPU network. It transforms raw text into highly descriptive 1024-dimensional vector embeddings and natively supports an expansive 8,192-token context window. Embeddings are indexed in Cloudflare Vectorize (humem-semantic-index) and coupled with a zero-latency SQLite database inside Durable Objects.
The Four-Layer Memory Model
Enterprise buyers evaluate memory architectures based on functional cognitive capabilities rather than raw database technologies. HuMem implements this through a standardized four-layer model:
1. The Soul File (Core)
Identity, Constraints & Beliefs. Represents the foundational core of the agent's identity. Managed through isolated registries to remain unaffected by adversarial prompt injections.
2. The User File
Multi-Session Synthesized User Profiles. Continuously refines raw conversational data into structured user preferences across disparate historical sessions to guarantee the agent adapts to user workflows.
3. The Agents File (Tool)
Registries & Collaborative Context. Manages the operational, tool, and collaborative state of the agent ecosystem. Facilitates cross-agent shared cognition securely.
4. The Memory File
Episodic & Semantic Knowledge Graph. Serves as the underlying structural substrate that records raw semantic facts and episodic timelines. Prevents the loss of temporal context during high-velocity data ingestion.
Ebbinghaus Forgetting Curve and Pruning Engine
To solve unbounded memory growth and context window pollution, HuMem.cloud implements a biologically-inspired Ebbinghaus memory decay algorithm.
- Dynamic Retention Scoring: The "strength" of a memory trace declines exponentially over time unless reinforced through active recall (the Spacing Effect).
- Category-Specific Half-Lives: Memory traces are routed into four distinct categories with specialized baseline half-lives (e.g., Strategy ~38 days, Fact ~24 days, Assumption ~19 days, Failure ~11 days).
- Autonomous Pruning: A background
alarm()dreaming cycle scans for inactive nodes where the calculated strength decays below the threshold. These nodes are summarized into dense narrative context, and the raw granular records are evicted from SQLite and Vectorize.