This is a more technical part of https://rkochanowski.com/article/embedding-benchmark/ sharing source code allowing you to verify and run it yourself. It contains all code snippets and detailed descriptions of all cases.
This benchmark evaluates how well an embedding model detects duplicated code. It focuses on code that does the same thing but is written differently. This is similarity detection, not retrieval.
Models are evaluated on how good they are at separating duplicated code from not duplicated, including adversarial cases.
- High-similarity cases focus on the same behavior, e.g., renamed identifiers, reordered statements, refactored structure, a different algorithm for the same problem, cross-language ports, a plain-English description of what the code does.
- Low-similarity cases focus on different behavior, e.g., shared identifiers, a small shared fragment, the same framework boilerplate, unrelated text.
config.yaml: Configuration of models, cases, and comparisons between two related cases (case_diff).env.example: Template of.envfile containing environment variables for API keys for providers used in benchmark
Gets embeddings and calculates cosine similarity for case pairs. Writes result/similarities.csv containing similarities for every pair of every case for every model.
uv run bench embedIt supports resume. When interrupted, you can safely re-run to continue. You can also add a new model keeping all previous data unchanged. Models are uniquely identified by label property.
Reads similarities.csv and generates gaps.csv, averages.csv, deviations.csv. The order of cases and models is determined by config.yaml.
uv run bench reportEach case directory contains case.yaml with a list of pairs and case description. Each pair has two files: side a and b as referenced in descriptions.
- exact-copy - Exact copies.
- formatting - Exact copies with different formatting, whitespace, or indentation.
- lang-similar-syntax - Same implementation in languages with similar syntax.
- lang-different-syntax - Same implementation in languages with different syntax.
- renamed-identifiers - Same implementation with different identifier names and the same literals.
- different-literals - Same implementation with the same identifiers and different literals.
- reordered-statements - Same logic with independent statements reordered.
- refactored-structure - Same logic with refactored structure.
- small-addition - Same code with one small addition.
- small-additions-large - Larger files with the same code plus several small additions.
- different-algorithm - Same problem solved with a different algorithm.
- transpiled - TypeScript and its transpiled JavaScript output.
- code-and-text - Code and text describing it.
- unrelated - Unrelated code in the same language with nothing else shared.
- same-identifiers - Unrelated code with heavily overlapping identifiers.
- small-shared-fragment - Different code with one small shared fragment.
- small-shared-fragments-large - Larger files with different code and several small shared fragments.
- same-boilerplate - Same boilerplate or scaffolding with different core logic.
- code-and-text - Code and unrelated text.