Skip to content

Load a model from memory - #4

Merged
mudler merged 1 commit into
mainfrom
feat/load-from-memory
Oct 4, 2026
Merged

mudler merged 1 commit into
mainfrom
feat/load-from-memory

Conversation

@localai-org-maint-bot

Copy link
Copy Markdown
Contributor

What

ced.cpp can now load a model from bytes in memory, not only from a file path. This helps a caller that already holds the model (a network blob, an archive, or one component of a bundle GGUF that packs several models) and does not want to write a temporary file or use a Linux-only in-memory file.

API

// Complete GGUF file in memory.
ced_ctx* ced_capi_load_from_memory(const void* data, size_t size);

// One model stored inside a larger GGUF: every key and tensor name is "<prefix><name>".
ced_ctx* ced_capi_load_from_memory_prefixed(const void* data, size_t size, const char* prefix);

C++: bool ced::Ced::load_from_memory(const void* data, size_t size, const std::string& prefix = "").

Ownership: the buffer is only read during the call. The tensor data is copied into the loader's own memory, so the caller may free or overwrite the buffer when the call returns, whatever the result. Peak memory during the call is about the buffer plus the model. The prefixed variant copies only the tensors under the prefix, so a bundle reader can pass the whole bundle and needs no standalone copy of the component. Errors follow the existing style: NULL, with the reason in ced_capi_last_error(NULL) (thread-local). ced_capi_abi_version() stays 1 because the change is additive. The path loader is unchanged.

How

  • The path loader and both memory loaders share everything after the I/O step (ModelLoader::read_model).
  • Parsing uses ggml's gguf_init_from_buffer, which is already in the pinned ggml. There is no temporary file or file descriptor, so the code is the same on Linux, macOS and Windows.
  • The prefixed loader parses only the header and tensor table, checks every tensor range against the buffer size, and copies the selected tensors into its own context under their unprefixed names.
  • The buffer is untrusted input, so it is checked first. ggml's reader aborts the process on some malformed input (for example a metadata key with an empty name), so a small precheck (src/gguf_check.cpp) walks the header and metadata with bounds checks before ggml sees the data. Metadata values are read only after their type is checked, because ggml also aborts on a type mismatch.

Tests

New ctest memory (tests/test_memory.cpp). It needs two different local CED models (tiny and mini Q8_0). It builds a two-component bundle in memory, and checks:

  • path, memory and prefixed loads (both component orders) give bitwise identical class scores;
  • the buffer is wiped and freed before classifying, so any kept pointer would show up under ASan;
  • NULL, empty, wrong-prefix and wrong-type input give a clear error;
  • truncation at about 1400 offsets for both a plain file and a bundle, and corrupt headers (bad magic, bad version, huge counts, all zero, all 0xff), are rejected;
  • a seeded bit-flip pass over the header (400 cases) never crashes;
  • 6 threads load from memory at the same time and still match.

Results on Linux x86-64: full ctest (8 tests, CPU, GGML_NATIVE=OFF Release) passes. The memory and capi tests pass under ASan + UBSan (with leak checking). The CLI smoke check from CI passes. Valgrind was not available.

Platforms

Only Linux was run. The new code is standard C++17 with no OS calls, and ggml's gguf_init_from_buffer is portable code, so macOS and Windows should behave the same, but I did not run them.

Limits

  • The check is on the file structure and the tensor ranges. A buffer that is well formed but holds a different or damaged model can still fail later (for example at classify time); this is the same as with a damaged file.
  • The unprefixed variant holds the buffer and the model in memory at once during the call.

🤖 Generated with Claude Code

ced.cpp could only load a model from a file path. A caller that already
holds the bytes (a network blob, an archive, or a bundle GGUF with
several models) had to write them to a file first.

Add two C functions and a C++ method:

* ced_capi_load_from_memory(data, size) loads a complete GGUF held in
  memory. The tensor data is copied during the call, so the caller can
  free the buffer when it returns.
* ced_capi_load_from_memory_prefixed(data, size, prefix) loads a model
  that is stored inside a larger GGUF, with every key and tensor name
  prefixed. Only the tensors under the prefix are copied, so a bundle
  reader does not need a standalone copy of the component.
* ced::Ced::load_from_memory(data, size, prefix = "").

The path loader and the memory loaders share everything after the I/O
step. Parsing uses ggml's gguf_init_from_buffer, so no temporary file or
file descriptor is needed and the code is the same on every platform.

The buffer comes from the caller, so it is checked before use. A new
precheck walks the header and metadata and rejects input that would
make ggml abort the process (for example an empty key name). Metadata
values are read only after their type is checked, and every tensor
range of a prefixed load is checked against the buffer size. Bad input
returns NULL with the reason in ced_capi_last_error(NULL).

test_memory compares path, memory and prefixed loads bit for bit, frees
the buffer before classifying, rejects truncated and corrupt buffers at
many offsets, runs a seeded bit-flip pass over the header, and loads
from several threads. It is clean under ASan and UBSan.

Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
@mudler
mudler merged commit 736a4ee into main Oct 4, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants