A small model that remembers what it sees.

YOLOX detects, SigLIP2 encodes identity, ManifoldFlow stores persistent object basins, and Gemma reads retrieved state through continuous latent tokens.

Research prototype · transient sessions
⚠️ Research prototype, not a validated perception system. The scripted scenarios below (car/box, cup/ball) are pre-recorded and always render correctly — they demonstrate the memory mechanism, not live detection. Live camera mode is a different story: the underlying detector was trained on a narrow, mostly-synthetic dataset and frequently misses or misclassifies real objects, including the trained toy car and box themselves. Treat every live-camera answer as a demo of the persistence mechanism working (or not) on top of imperfect, sometimes-wrong perception — not as evidence of general-purpose object recognition.
Starting private session…

Persistent state — what ManifoldFlow currently remembers

No entities remembered yet.

Observed events

Camera access is opt-in. Frames are resized, processed in memory, and are not retained. No microphone is requested. Sessions expire after ten minutes.

Ask the world

visual — answered from what the camera has seenrecalled — answered from something you said earlier this session
Try “Where is the car?”, “Which cup contains the ball?”, or “What happened?”

Retrieved evidence

No query yet.