Multi-Modal AI

AI Memory: multimodal semantic search for AI agents

By Khadim Hussain · · 1 min read

In short

A self-hosted memory service for AI agents. Text, images, audio and video go into one embedding space (CLIP ViT-H-14 and Whisper large-v3), Qdrant runs the similarity search, and a PostgreSQL graph links related memories. Memories fade over time unless an agent uses them again.

4
modalities
<50ms
latency
6+
services
  • Python
  • FastAPI
  • Qdrant
  • CLIP
  • Whisper
  • PostgreSQL
  • Apache AGE
  • Redis
  • Celery
  • MinIO
  • Docker
  • CUDA
AI Memory system architecture and API
On this page

The problem

AI agents lose context between sessions, and most memory layers only handle text. An agent that has seen a diagram, a voice note or a screen recording has no way to find it again by meaning.

How it works

The system is a set of small services behind one REST API:

LayerToolJob
APIFastAPIMemory CRUD, semantic search, reinforcement and association endpoints, with authentication, rate limiting and tenant isolation
EncodersCLIP ViT-H-14, Whisper large-v3 on CUDAEmbed images and text; transcribe audio and video so they land in the same vector space
Vector searchQdrantSimilarity search across all modalities
AssociationsPostgreSQL with Apache AGEA graph of links between related memories
FilesMinIOS3-compatible storage for the original media, isolated per tenant
Background jobsRedis and CeleryIngestion and batch encoding off the request path

Because every modality ends up in one embedding space, a text query can return an image or a clip of audio, and the other way round.

Memory that fades

Stored memories follow a lifecycle modeled on how people remember. Each memory decays along an Ebbinghaus forgetting curve, gets reinforced when an agent retrieves it, and moves from short-term to long-term storage once it has been reinforced enough. Agents keep what they use and lose what they never touch.

Where this fits

The service is meant as self-hosted infrastructure for RAG over mixed media: support bots that need screenshots and call recordings, research assistants over papers and talks, or any agent that should remember what it was shown.

Want a second pair of eyes on this?

Book a 15-minute intro call. Bring the problem, and you leave with a concrete next step.