module
Memo::Storage
Overview
Low-level storage operations for embeddings and chunks
Extended Modules
Defined in:
memo/storage.crInstance Method Summary
-
#compute_hash(text : String) : Bytes
Compute SHA256 hash for text content
-
#create_chunk(db : DBHandle, hash : Bytes, source_type : String, source_id : Int64, offset : Int32 | Nil, size : Int32, pair_id : Int64 | Nil = nil, parent_id : Int64 | Nil = nil) : Int64
Create chunk reference (or ignore if already exists)
-
#decode_vector(blob : Bytes) : Array(Float64)
Decode a vector stored by encode_vector.
-
#deserialize_embedding(blob : Bytes) : Array(Float64)
Deserialize embedding from binary blob
-
#encode_vector(vector : Array(Float64)) : Bytes
Encode a vector for memo_vectors as little-endian IEEE half floats.
-
#get_rowid(db : DBHandle, hash : Bytes, service_id : Int64) : Int64 | Nil
Get the rowid of an embedding by hash and service_id.
-
#get_service_by_format_model(db : DBHandle, format : String, model : String) : Tuple(Int64, String, String | Nil, String, Int32, Int32, Float64) | Nil
Returns service record by format and model, or nil if not found
-
#get_service_by_name(db : DBHandle, name : String) : Tuple(Int64, String, String | Nil, String, Int32, Int32, Float64) | Nil
Get service by name
-
#increment_match_count(db : DBHandle, chunk_ids : Array(Int64))
Increment match_count for chunks
-
#increment_read_count(db : DBHandle, chunk_ids : Array(Int64))
Increment read_count for chunks
-
#register_service(db : DBHandle, name : String | Nil, format : String, base_url : String | Nil, model : String, dimensions : Int32, max_tokens : Int32) : Int64
Register or get existing service by name
-
#serialize_embedding(embedding : Array(Float64)) : Bytes
Serialize embedding to binary blob (Int16 for 50% storage reduction)
-
#store_embedding(db : DBHandle, hash : Bytes, token_count : Int32, service_id : Int64) : Tuple(Bool, Int64)
Register embedding hash in database (deduplicated by hash + service_id)
-
#update_tokens_per_byte(db : DBHandle, service_id : Int64, observed_ratio : Float64)
Update tokens_per_byte ratio using exponential moving average
Instance Method Detail
Create chunk reference (or ignore if already exists)
Returns chunk id if inserted, or 0 if chunk already existed (was ignored)
Encode a vector for memo_vectors as little-endian IEEE half floats.
That is the precision the USearch index keeps (f16 quantization), so an index rebuilt from stored vectors matches one built from the provider's. Unlike serialize_embedding, values outside -1..1 aren't clamped.
Get the rowid of an embedding by hash and service_id.
Returns service record by format and model, or nil if not found
Get service by name
Increment match_count for chunks
Increment read_count for chunks
Register or get existing service by name
Serialize embedding to binary blob (Int16 for 50% storage reduction)
Register embedding hash in database (deduplicated by hash + service_id)
Returns {inserted, rowid} where inserted is true if new, rowid is the USearch key.
Update tokens_per_byte ratio using exponential moving average