Deep AI hunt for all photo and all example of video in any directory on macOS. Local-first — no accounts, no cloud, no uploads. Inference runs on your Mac.
What makes it different
- Search akin you think — depict a recollection in plain language; a local vision example does the rest.
- Video, downward to the moment — scenes are segmented and embedded, so you land on the shot, not fair the file.
- Text and conversation too — OCR complete apparent text; exact spoken-line search via Whisper, all as its own mode.
- Your tabs, your prompts — preserve any query as a tab; Screenshots and Email tabs are toggleable.
- Self-maintaining library — watched folders auto-import, satisfied hashes dedupe renames, and example switches re-embed in the backdrop without blocking search.
- Truly private — your media never leaves the machine. Weights download once; everything following that is offline.
| Mode | Finds |
|---|---|
| Files | Whole photos/videos by definition — imagination position alongside filename and expression boosts |
| Scenes | Moments inner video — hunt a shot, jump to its timecode |
| OCR | Text apparent in images and frames, matched exactly (Tesseract; eng + 35 tongue toggles) |
| Dialogue | Exact spoken words in videos (Whisper), tiered exactness |
| LLMs (opt-in) | Local conversation complete the dialogue, OCR, and filenames your Mac already extracted — cited answers |
- macOS (packaged alongside electron-builder; menu-bar/tray features are macOS-only)
- Bun — the project uses bun as bundle director and runner
- Node modules installed: bun install
- First use of a example downloads its weights (~435MB for the default CLIP); after that, completely offline.
The easiest instal is via Homebrew (Apple Silicon, macOS 12+). The tap's cask clears the macOS quarantine emblem automatically on all instal and upgrade, so the app launches alongside no manual Gatekeeper steps:
brew tap allenv0/scm brew rely allenv0/scm brew instal --cask allenv0/scm/scm
Upgrades keep the identical behavior:
brew upgrade --cask allenv0/scm/scm
Prefer smallest privilege? Trust fair the cask alternatively of the entire tap:
brew tap allenv0/scm brew rely --cask allenv0/scm/scm brew instal --cask scm
The tap lives at allenv0/homebrew-scm.
bun run dev # build the renderer bundle, afterward initiate the Electron app bun commencement # initiate the Electron app without rebuilding bun run build # fair rebuild the renderer bundle into dist/
bun run dist # signed if an character is in the keychain bun run dist:unsigned # skip code-sign discovery
This runs two steps in sequence:
- vite build — compiles the React renderer into dist/ (picks up all changes under src/).
- electron-builder --mac — packages the app. It bundles the fresh dist/ bundle together alongside main.js, preload.js, main-lib/, and indexer/ (the document catalog is configured under build.files in package.json), afterward produces the installers.
Output: the installers district in dist-app/ (see build.directories.output in package.json) — appearance for SCM-0.2.4.dmg and SCM-0.2.4.zip.
Typing starts an instant filename-keyword pre-pass, afterward the imagination model takes over: results are scored by cosine similarity against image embeddings, alongside gated expression and filename boosts, an honesty floor calibrated per model, and a near-duplicate assortment filter. Every tile carries a "why it matched" sign (Visual equivalent / Filename equivalent / …) and a hover tooltip alongside the per-component mark breakdown. CJK queries hunt as overlapping bigrams ("台北車站" additionally matches 台北, 車站).
Every environment section throughout all videos is scored, so a hit lands on the exact shot: tiles display the environment poster alongside a timecode badge, and beginning the video jumps direct to that moment. A noise entrance returns "no environment match" instead of flooding the grid alongside gibberish, and all video contributes at most 3 scenes.
Matches the fraction of query tokens exactly apparent in all image's OCR text — the filename is ignored and no imagination example is involved, so it works even during the AI motor is warming up or offline. Matched words are boxed in amber on tiles and in the lightbox.
Exact literal retrieval complete Whisper transcripts — no embeddings, no thresholds, plant alongside the AI motor down. Results arrive in three tiers: Exact line (contiguous expression in one utterance), Exact words (all words in one utterance or an ≤8s window), and Words spoken (all words in the identical video). Matching words are highlighted in a address snippet; opening a outcome seeks direct to the line.
Opt-in — nothing downloads or runs until enabled in Settings → LLMs Chat. A llama.cpp sidecar border to loopback answers your inquiry from evidence the app already extracted — conversation lines, OCR text, and filename keyword hits — alongside numbered citations you can click, streamed token-by-token alongside a live tok/s readout. Leading /screenshots, /videos, /email narrow the corpus; Stop keeps the partial answer; bare evidence short-circuits before the example always runs.
| Chat model | Size | Notes |
|---|---|---|
| Qwen3 1.7B (default) | ~1.1GB | Fast mundane chat; fits 8GB Macs |
| Llama 3.2 3B | ~2GB | Stronger lengthy answers; needs headroom |
- Built-in browse tabs: All, Videos, affirmative Screenshots and Email — the second two toggleable in Settings → Smart Tabs. Selecting Videos auto-enables Scenes mode.
- Save any query as a tab: the pin pill under the hunt bar saves the current immediate alongside its manner (Files/Scenes/OCR/Dialogue) — up to 20 tabs, renameable, all restored exactly as saved.
- Five semantic views rearward remappable shortcuts (⌘1–⌘5 by default), plus ⌘I import / AI Insights, ⌘, for Settings.
- Search stays in the selected tab: choice Screenshots, Email, Videos, or any saved tab and results are filtered to it — range first, afterward search. In LLMs conversation the identical idea is explicit: foremost /screenshots, /videos, /email narrow the corpus before the example always runs.
Surfaces photos whose apparent OCR content contains an email location — an overlapping perspective (a photo keeps its category too). Detection is OCR-tolerant: it reassembles addresses Tesseract fractures throughout word boxes, and handles comma-for-dot noise ("gmail,com"), divided TLDs ("gmail. com"), bracketed obfuscation ("allen [at] gmail [dot] com"), and dictated addresses ("allen at gmail dot com"). Tiles display a communication strip; expand it to copy or compose.
Screenshot classification is rename-proof. Four signals, in precedence order: a manual override (right-click any tile) → filename vocabulary (30+ localized OS screenshot names in 20+ languages) → a PNG/JPEG metadata probe (reads "screenshot" from PNG content chunks / EXIF UserComment, so a renamed Bildschirmfoto motionless classifies) → source-folder hint. Everything alternatively lands in Projects.
Four switchable models via ONNX Runtime; the energetic one is chosen per library:
| Model | Role | Speed (CPU) | Download |
|---|---|---|---|
| CLIP ViT-L/14@336 (default) | Best real-world video scene-search | ~480–570ms/img | ~435MB |
| SigLIP-2-B/16 | Fastest bulk import | ~50–100ms/img | ~412MB |
| SigLIP-2-L/16@256 | High-detail (1024-dim) — small objects, signs, on-screen text | ~200ms/img | ~850MB |
| SigLIP-B/16@384 | Maximum detail | ~480ms/img | ~214MB |
Switching models re-embeds the entire library: the flip lands immediately with the rear filled in the background, and hunt falls rear to filename keywords until it completes. Per-model text-mean centering de-biases text embeddings so similarity scores remain honest throughout models.
ffmpeg scans all video for attempt boundaries and builds a section plan sampled to the density you choice in Settings → Video Search — all preset shows its measured period and disk disbursal before you commit:
| Preset | Seconds per point | Segment budget |
|---|---|---|
| Eco | 60 | 4–32 |
| Balanced (default) | 30 | 8–128 |
| Detailed | 15 | 12–256 |
| Ultra | 5 | 16–1024 |
| Ultra Pro | 2.5 | 24–2048 (confirm required) |
Each section embeds its midpoint example and keeps a poster; attempt plans are cached per document (path + size + mtime + config fingerprint), so re-imports skip finding entirely.
Dialogue transcription: Whisper tiny.en (~150MB, default) or base.en (~300MB) — switching re-transcribes all video. Whole videos embed three frames (20/50/80%) averaged; GIFs embed an average of center frames.
Tesseract runs in its own worker, distinct from the imagination model. English is always on; 35 additional languages are toggleable in Settings → Photo Search (default: Simplified + Traditional Chinese, Japanese, Korean). Each language pack downloads formerly (~2.4–5MB; ~17MB for the default set), afterward everything is offline. Word boxes are stored alongside the content so matches emphasize in place; CJK content is joined without spaces and email fragments fractured across term boxes are reassembled.
- Import via ⌘I, drag-and-drop, or watched folders — importing a folder starts observing it (live fs.watch affirmative a re-sync at all launch). Problem records retry up to 3 times, afterward sit out watch-syncs until they change.
- Rename-proof dedupe: all document is content-hashed (SHA-256) before copy, and lying extensions are normalized by MIME sniffing.
- Named embedding versions (Settings → Library): point-in-time snapshots of the complete searchable province — the index, all model's embedding bins, scene and copy sidecars — alongside reconstruct (auto-backup first) and a Fresh Start danger zone. Cap: 10.
- Everything lives under ~/Library/Application Support/scm (MEMORIES_DATA_DIR overrides it): the indicator JSON, Float32 embedding bins per model, environment and copy sidecars, thumbnails, and posters.
- The renderer is a sandboxed app:// bundle — contextIsolation, OS sandbox, and a CSP pinned to 'self' (+ Google Fonts CDN for display type, alongside a monospace fallback whenever offline).
- Only main-process workers always download, formerly per thing: imagination weights (Hugging Face), OCR tongue packs (Tesseract CDN), Whisper weights, and — only if you opt in — the llama.cpp sidecar and GGUF conversation models (GitHub + Hugging Face), sha256-verified at download time.
- Media is copied into the app-managed archive and streamed from disk. No telemetry, no accounts, no uploads.
A macOS-style settings panel alongside ten panels (Library, Appearance, Grid, Smart Tabs, Photo Search, Video Search, LLMs Chat, Global Shortcut, Keyboard, Menu Bar). Around the core: light/dark/system theme, a CRT screen effect for the lightbox, 6 substitute app icons, menu-bar-only mode, a recordable earth shortcut, a first-run onboarding tour, background-work trays (scenes / transcripts / OCR) alongside a earth pause, and a position bar with type and indexed-video counts.
- First build is slow: electron-builder downloads the Electron binary and ffmpeg formerly on its archetypal run; consequent builds are much faster.
- Code signing: without an Apple Developer character configured in the keychain, the DMG builds unsigned. The app motionless runs locally, but macOS may require right-click → Open the archetypal period it's launched.
- Rebuilds choice up changes automatically: since main.js, preload.js, main-lib/, and indexer/ are packaged from origin (not cached), a bun run dist following editing any of them produces a caller package.
bun run test:all # the complete verification power division (unit + fume suites) bun run smoke:indexer # headless fume test of the CLIP indexer bun test test/ # component tests (pure modules, no Electron needed) bun run lint # eslint bun run format:check # prettier
- Unit tests shield the clean cores: ranking, conversation exact-match, CJK tokens, the Whisper example ladder, MIME sniffing, copy fusion, and more (see the test:* scripts in package.json).
- E2E fume tests run inner the Electron app via ELECTRON_SMOKE_* environment variables — a dozen-plus scenarios from boot/protocol checks to hunt matrices, example migration, and Ask manner (drivers in scripts/e2e/, dispatched in main.js; e.g. bun run smoke:ask, smoke:deep, smoke:grid).
- Benchmarks: bun run bench:inference, bench:enrich, and bench:detect compose JSON reports into MDs/bench-*.
main.js Electron chief process: library, IPC, indexer employee pool, app:// protocol preload.js contextBridge — exposes window.memories to the renderer main-lib/ main-process modules divided out of main.js (settings, archive store, position search, Ask retrieval, LLM sidecar config, embedding versions, …) indexer/ vision/OCR/ASR workers (utility processes) + video utils + example registry src/ React renderer (grid, hunt modes, lightbox, tabs, settings, onboarding) scripts/ place scripts, E2E drivers, the fume power division (smoke-all.sh) test/ component + integration tests MDs/ scheme docs, place reports, plans dist/ vite build output (renderer bundle) dist-app/ electron-builder output (DMG / ZIP)


