Repository navigation
Load a model from memory - #4
Merged
Merged
Conversation
ced.cpp could only load a model from a file path. A caller that already holds the bytes (a network blob, an archive, or a bundle GGUF with several models) had to write them to a file first. Add two C functions and a C++ method: * ced_capi_load_from_memory(data, size) loads a complete GGUF held in memory. The tensor data is copied during the call, so the caller can free the buffer when it returns. * ced_capi_load_from_memory_prefixed(data, size, prefix) loads a model that is stored inside a larger GGUF, with every key and tensor name prefixed. Only the tensors under the prefix are copied, so a bundle reader does not need a standalone copy of the component. * ced::Ced::load_from_memory(data, size, prefix = ""). The path loader and the memory loaders share everything after the I/O step. Parsing uses ggml's gguf_init_from_buffer, so no temporary file or file descriptor is needed and the code is the same on every platform. The buffer comes from the caller, so it is checked before use. A new precheck walks the header and metadata and rejects input that would make ggml abort the process (for example an empty key name). Metadata values are read only after their type is checked, and every tensor range of a prefixed load is checked against the buffer size. Bad input returns NULL with the reason in ced_capi_last_error(NULL). test_memory compares path, memory and prefixed loads bit for bit, frees the buffer before classifying, rejects truncated and corrupt buffers at many offsets, runs a seeded bit-flip pass over the header, and loads from several threads. It is clean under ASan and UBSan. Assisted-by: Claude:claude-sonnet-5-5 [Claude Code]
mudler
approved these changes
Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
ced.cpp can now load a model from bytes in memory, not only from a file path. This helps a caller that already holds the model (a network blob, an archive, or one component of a bundle GGUF that packs several models) and does not want to write a temporary file or use a Linux-only in-memory file.
API
C++:
bool ced::Ced::load_from_memory(const void* data, size_t size, const std::string& prefix = "").Ownership: the buffer is only read during the call. The tensor data is copied into the loader's own memory, so the caller may free or overwrite the buffer when the call returns, whatever the result. Peak memory during the call is about the buffer plus the model. The prefixed variant copies only the tensors under the prefix, so a bundle reader can pass the whole bundle and needs no standalone copy of the component. Errors follow the existing style:
NULL, with the reason inced_capi_last_error(NULL)(thread-local).ced_capi_abi_version()stays 1 because the change is additive. The path loader is unchanged.How
ModelLoader::read_model).gguf_init_from_buffer, which is already in the pinned ggml. There is no temporary file or file descriptor, so the code is the same on Linux, macOS and Windows.src/gguf_check.cpp) walks the header and metadata with bounds checks before ggml sees the data. Metadata values are read only after their type is checked, because ggml also aborts on a type mismatch.Tests
New ctest
memory(tests/test_memory.cpp). It needs two different local CED models (tiny and mini Q8_0). It builds a two-component bundle in memory, and checks:Results on Linux x86-64: full ctest (8 tests, CPU,
GGML_NATIVE=OFFRelease) passes. Thememoryandcapitests pass under ASan + UBSan (with leak checking). The CLI smoke check from CI passes. Valgrind was not available.Platforms
Only Linux was run. The new code is standard C++17 with no OS calls, and ggml's
gguf_init_from_bufferis portable code, so macOS and Windows should behave the same, but I did not run them.Limits
🤖 Generated with Claude Code