module Memo::IndexJournal

Overview

Keeps a service's USearch index recoverable from the database.

Every vector memo stores goes into memo_vectors, and every embedding added or removed is logged in memo_index_log, in the same transaction as the change. Saving the index records a checkpoint beside the file (the last log entry it includes) and prunes the log up to it.

Opening the index replays only the entries after the checkpoint, so a clean restart does no work. When the file is missing, corrupt, or older than the pruned log, the index is rebuilt from memo_vectors. Neither path calls the embedding API.

Replay is idempotent: each logged embedding is set to its current state in the database (added if it has a stored vector, removed otherwise). A checkpoint that lags the file only costs extra replay.

Extended Modules

Defined in:

memo/index_journal.cr

Constant Summary

REBUILD_PAGE = 256

Vectors read per query during a rebuild. Reading in pages keeps no cursor open while other fibers run.

YIELD_EVERY = 16

Replay and rebuild hand the thread to other fibers every this many vectors, so opening an index doesn't stall the rest of a server (an insert takes ~4 ms at 1536 dimensions).

Instance Method Summary

Instance Method Detail

def catch_up(db : DB::Database, index : USearch::Index, service_id : Int64, path : String) : Nil #

Bring an open index back in step with the database after a failed update left committed changes out of it: replay the log since the file's checkpoint (or rebuild, if that part is already pruned). Replay is idempotent, so entries the index already has do no harm.


[View source]
def checkpoint(db : DB::Database, index : USearch::Index, service_id : Int64, path : String) : Nil #

Save the index and checkpoint the journal: record the last log entry the saved file includes, then prune the log up to it. The file is written on a separate thread (see USearchIndex.save_in_background).

The caller must make sure every committed change has been applied to index, and keep writers out until this returns, or the checkpoint would claim changes the file lacks.


[View source]
def open(db : DB::Database, path : String, dimensions : Int32, service_id : Int64) : Tuple(USearch::Index, Recovery, File) #

Lock and open the index at path, bringing it up to date with the database. Returns the index, what recovery did, and the lock to hold while the index is open (see USearchIndex.lock).


[View source]
def record_removals(cnn : DB::Connection, pending : USearchIndex::Pending, hash : Bytes, service_id : Int64) : Nil #

Record the removal of every service's embedding of hash, inside the caller's transaction. Only service_id's index is open here; other services pick up their removals from the log when they next open.

Limitation: if another Service has one of those indexes open at the same time, its next save can checkpoint past the removal, leaving a stale vector in that index. Searches filter it out (its embedding row is gone), but it takes up one of the nearest-neighbor slots.


[View source]
def record_vector(cnn : DB::Connection, pending : USearchIndex::Pending, embedding_id : Int64, service_id : Int64, vector : Array(Float64), inserted : Bool) : Nil #

Store a vector for an embedding row, inside the caller's transaction, and queue it for the index. A row that already has a stored vector (deduplicated content) is skipped; an older row without one gets it.


[View source]