module
Memo::IndexJournal
Overview
Keeps a service's USearch index recoverable from the database.
Every vector memo stores goes into memo_vectors, and every embedding added or removed is logged in memo_index_log, in the same transaction as the change. Saving the index records a checkpoint beside the file (the last log entry it includes) and prunes the log up to it.
Opening the index replays only the entries after the checkpoint, so a clean restart does no work. When the file is missing, corrupt, or older than the pruned log, the index is rebuilt from memo_vectors. Neither path calls the embedding API.
Replay is idempotent: each logged embedding is set to its current state in the database (added if it has a stored vector, removed otherwise). A checkpoint that lags the file only costs extra replay.
Extended Modules
Defined in:
memo/index_journal.crConstant Summary
-
REBUILD_PAGE =
256 -
Vectors read per query during a rebuild. Reading in pages keeps no cursor open while other fibers run.
-
YIELD_EVERY =
16 -
Replay and rebuild hand the thread to other fibers every this many vectors, so opening an index doesn't stall the rest of a server (an insert takes ~4 ms at 1536 dimensions).
Instance Method Summary
-
#catch_up(db : DB::Database, index : USearch::Index, service_id : Int64, path : String) : Nil
Bring an open index back in step with the database after a failed update left committed changes out of it: replay the log since the file's checkpoint (or rebuild, if that part is already pruned).
-
#checkpoint(db : DB::Database, index : USearch::Index, service_id : Int64, path : String) : Nil
Save the index and checkpoint the journal: record the last log entry the saved file includes, then prune the log up to it.
-
#open(db : DB::Database, path : String, dimensions : Int32, service_id : Int64) : Tuple(USearch::Index, Recovery, File)
Lock and open the index at
path, bringing it up to date with the database. -
#record_removals(cnn : DB::Connection, pending : USearchIndex::Pending, hash : Bytes, service_id : Int64) : Nil
Record the removal of every service's embedding of
hash, inside the caller's transaction. -
#record_vector(cnn : DB::Connection, pending : USearchIndex::Pending, embedding_id : Int64, service_id : Int64, vector : Array(Float64), inserted : Bool) : Nil
Store a vector for an embedding row, inside the caller's transaction, and queue it for the index.
Instance Method Detail
Bring an open index back in step with the database after a failed update left committed changes out of it: replay the log since the file's checkpoint (or rebuild, if that part is already pruned). Replay is idempotent, so entries the index already has do no harm.
Save the index and checkpoint the journal: record the last log entry the saved file includes, then prune the log up to it. The file is written on a separate thread (see USearchIndex.save_in_background).
The caller must make sure every committed change has been applied to
index, and keep writers out until this returns, or the checkpoint
would claim changes the file lacks.
Lock and open the index at path, bringing it up to date with the
database. Returns the index, what recovery did, and the lock to hold
while the index is open (see USearchIndex.lock).
Record the removal of every service's embedding of hash, inside the
caller's transaction. Only service_id's index is open here; other
services pick up their removals from the log when they next open.
Limitation: if another Service has one of those indexes open at the same time, its next save can checkpoint past the removal, leaving a stale vector in that index. Searches filter it out (its embedding row is gone), but it takes up one of the nearest-neighbor slots.
Store a vector for an embedding row, inside the caller's transaction, and queue it for the index. A row that already has a stored vector (deduplicated content) is skipped; an older row without one gets it.