The problem
AI agents lose context between sessions, and most memory layers only handle text. An agent that has seen a diagram, a voice note or a screen recording has no way to find it again by meaning.
How it works
The system is a set of small services behind one REST API:
| Layer | Tool | Job |
|---|---|---|
| API | FastAPI | Memory CRUD, semantic search, reinforcement and association endpoints, with authentication, rate limiting and tenant isolation |
| Encoders | CLIP ViT-H-14, Whisper large-v3 on CUDA | Embed images and text; transcribe audio and video so they land in the same vector space |
| Vector search | Qdrant | Similarity search across all modalities |
| Associations | PostgreSQL with Apache AGE | A graph of links between related memories |
| Files | MinIO | S3-compatible storage for the original media, isolated per tenant |
| Background jobs | Redis and Celery | Ingestion and batch encoding off the request path |
Because every modality ends up in one embedding space, a text query can return an image or a clip of audio, and the other way round.
Memory that fades
Stored memories follow a lifecycle modeled on how people remember. Each memory decays along an Ebbinghaus forgetting curve, gets reinforced when an agent retrieves it, and moves from short-term to long-term storage once it has been reinforced enough. Agents keep what they use and lose what they never touch.
Where this fits
The service is meant as self-hosted infrastructure for RAG over mixed media: support bots that need screenshots and call recordings, research assistants over papers and talks, or any agent that should remember what it was shown.
