Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 25 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -6,5 +6,29 @@ on:
jobs:
label-gate:
uses: AustralianCancerDataNetwork/cava-devops/.github/workflows/label-gate.yml@main
build-test:
build-test-sqlite:
uses: AustralianCancerDataNetwork/cava-devops/.github/workflows/build-test.yml@main
build-test-postgres:
uses: AustralianCancerDataNetwork/cava-devops/.github/workflows/build-test-postgres-v2.yml@main
with:
postgres-db: omop_graph_test
setup-commands: |
uv run omop-config configure omop_graph \
--set cdm_db.kind=cdm \
--set cdm_db.connection.dialect=postgresql+psycopg \
--set cdm_db.connection.host=localhost \
--set cdm_db.connection.port=5432 \
--set cdm_db.connection.user=test \
--set cdm_db.connection.password=test \
--set cdm_db.connection.database_name=omop_graph_ci_placeholder \
--set cdm_db.connection.test_only=false \
--set cdm_db.schema_name=public \
--set test_cdm_db_pg.kind=cdm \
--set test_cdm_db_pg.connection.dialect=postgresql+psycopg \
--set test_cdm_db_pg.connection.host=localhost \
--set test_cdm_db_pg.connection.port=5432 \
--set test_cdm_db_pg.connection.user=test \
--set test_cdm_db_pg.connection.password=test \
--set test_cdm_db_pg.connection.database_name=omop_graph_test \
--set test_cdm_db_pg.connection.test_only=true \
--set test_cdm_db_pg.schema_name=public
3 changes: 0 additions & 3 deletions Dockerfile

This file was deleted.

4 changes: 2 additions & 2 deletions docs/graph/kg.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,7 +62,7 @@ This requires the optional `omop-emb` package — see the [installation guide](.
```python
from sqlalchemy import create_engine
from oa_configurator import Resolver
from omop_emb.backends import resolve_backend_from_resolved_vector_store
from omop_emb.backends import open_vector_store_writer
from omop_graph.graph.kg import KnowledgeGraph, KnowledgeGraphEmbeddingConfiguration
from omop_emb.config import MetricType

Expand All @@ -71,7 +71,7 @@ engine = create_engine("postgresql://user:pass@localhost/omop")
resolver = Resolver.from_active_config()
resolved_vector_store = resolver.resolve_vector_store("vector_store") # a [vector_stores.*] entry name
resolved_model = resolver.resolve_model("embedding-model") # a [models.*] entry name
backend = resolve_backend_from_resolved_vector_store(resolved_vector_store)
backend = open_vector_store_writer(resolved_vector_store)

emb_config = KnowledgeGraphEmbeddingConfiguration(
metric_type=MetricType.COSINE,
Expand Down
29 changes: 20 additions & 9 deletions docs/oaklib/interface.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,31 +29,42 @@ The primary adapter class that inherits from multiple OAK interfaces:
* **`TextAnnotatorInterface`**: Provides a pipeline to ground raw text spans to OMOP Concept IDs using the internal `KnowledgeGraph` resolvers.

### Resource Management
To initialize a connection, `omop-graph` uses a specialized resource factory:
The adapter resolves its database through oa-configurator; there is no URL-based construction.

* **`OMOPOntologyResource`**: A dataclass that wraps the SQLAlchemy connection URL, treating the database as a live ontology source.
* **`omop_resource()`**: A factory function that resolves database credentials from an explicit URL, or from the active oa-configurator stack config (`OmopGraphConfig.cdm_db`) when no URL is given.
* **`OMOPOntologyResource`**: The oaklib resource for the `omop` scheme. Its `slug` names a `[databases.*]` entry; no slug resolves `OmopGraphConfig.cdm_db`.
* **`resolved=`**: Pass an already-resolved `ResolvedCDMDatabase` to build the engines from it directly. A separate `vocab_connection` produces a second, vocabulary-only engine.
* **`kg=`**: Pass an existing `KnowledgeGraph` to use its engines as-is.

---

## Usage Examples

### Initializing the Adapter
You can get an OAK-compliant adapter by providing a SQLAlchemy connection string:
Select the adapter through oaklib with an oa-configurator database name:

```python
from omop_graph.oaklib_interface import OMOPAlchemyImplementation
from oaklib import get_adapter

# Initialize via connection string
adapter = OMOPAlchemyImplementation(
engine_string="postgresql://user:pass@localhost/omop"
)
# Named [databases.*] entry
adapter = get_adapter("omop:cdm_db")

# Or OmopGraphConfig.cdm_db
adapter = get_adapter("omop:")

# Use standard OAK methods
label = adapter.label("OMOP:44819488")
print(f"Label: {label}")
```

Or construct it from an already-resolved database:

```python
from omop_graph.db.session import resolve_cdm_database
from omop_graph.oaklib_interface.omop_implementation import OMOPAlchemyImplementation

adapter = OMOPAlchemyImplementation(resolved=resolve_cdm_database("cdm_db"))
```

### Searching and Traversal
Since the adapter implements the `SearchInterface`, you can perform standardized searches:

Expand Down
12 changes: 6 additions & 6 deletions docs/reasoning/grounding.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ To accelerate the grounding to standard concepts, `omop-graph` makes use of:
| `parent_ids` | `tuple[int, ...]` | `None` | Only accept candidates that are descendants of these OMOP concept IDs (hierarchy validation via `concept_ancestor`). |
| `search_constraint` | `ConceptFilter` | `None` | Filters applied to the initial resolver query (concept IDs, domain, vocabulary, standard/active flags, and limit). |
| `max_depth` | `int` | `6` | Maximum hop distance allowed between a candidate and its standard anchor. |
| `predicate_kinds` | `frozenset[PredicateKind]` | `{IDENTITY}` | Relationship kinds followed when walking from a non-standard candidate to its standard anchor. |
| `predicate_kinds` | `frozenset[PredicateKind]` | `{IDENTITY}` | Relationship kinds followed when walking from a non-standard candidate to its standard anchor. Currently locked: `__post_init__` raises `ValueError` for any value other than exactly `frozenset({PredicateKind.IDENTITY})` — not yet configurable in practice, despite being a normal dataclass field. |

### ConceptFilter

Expand Down Expand Up @@ -99,15 +99,15 @@ TotalScore = Relevance - ParsimonyPenalty + BroadnessBonus
$$

#### 1. Relevance
Relevance represents the initial semantic fit and is computed as **either** embedding similarity **or** textual similarity — not both simultaneously:
Relevance represents the initial semantic fit and is computed as **either** embedding similarity **or** textual similarity, chosen per candidate rather than globally — not both simultaneously for the same candidate:

- **Without embeddings**: textual similarity is used exclusively.
- **With embeddings** (default when `omop-graph[emb]` is installed and configured): embedding cosine similarity **replaces** the textual score entirely.
- A candidate resolved by `EmbeddingResolver` (`match_kind == LabelMatchKind.EMBEDDING`) gets embedding cosine similarity.
- Every other candidate — resolved via exact/partial/full-text matching — always gets textual similarity, whether or not `omop-graph[emb]` is installed.

The two scoring modes:

- **Embedding Similarity**: Cosine similarity between the input text embedding and the concept embedding. Requires `omop-graph[emb]` and a configured `KnowledgeGraphEmbeddingConfiguration` — see the [Knowledge Graph docs](../graph/kg.md#embedding-configuration) and the [omop-emb documentation](https://australiancancerdatanetwork.github.io/omop-emb/) for setup.
- **Textual Similarity**: A custom token-overlap score that heavily penalizes missing words from the user's query but allows for "extra" descriptive words in the OMOP concept name. Used as a fallback when no embedding is available.
- **Embedding Similarity**: Cosine similarity between the input text embedding and the concept embedding. Only applies to candidates resolved via `EmbeddingResolver`, which requires `omop-graph[emb]` and a configured `KnowledgeGraphEmbeddingConfiguration` — see the [Knowledge Graph docs](../graph/kg.md#embedding-configuration), [Resolver Pipelines](resolvers.md), and the [omop-emb documentation](https://australiancancerdatanetwork.github.io/omop-emb/) for setup.
- **Textual Similarity**: A custom token-overlap score that heavily penalizes missing words from the user's query but allows for "extra" descriptive words in the OMOP concept name. Used for every candidate not resolved via embeddings.

#### 2. Parsimony: Distance Penalty
OMOP is a deep hierarchy. A concept that is 1 hop away from your search term is more likely to be correct than one found 5 hops away.
Expand Down
7 changes: 4 additions & 3 deletions docs/reasoning/resolvers.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,14 +6,15 @@ The backbone of the `ResolverPipeline` are specific **resolvers**. `omop-graph`

- **`ExactLabelResolver`**: Exact case-insensitive match. Is there anywhere in the the [`concept`](https://ohdsi.github.io/CommonDataModel/cdm54.html#concept) table the exact string `"Hodgkin lymphoma"`
- **`ExactSynonymResolver`**: Similar to `ExactLabelResolver` yet searching the [`concept_synonym`](https://ohdsi.github.io/CommonDataModel/cdm54.html#concept_synonym) table
- **`FullTextResolver`**: Matches irrespective of word order. Not as relevant for the example above but relevant for others (e.g., "Kidney Cancer" -> "Cancer of Kidney").
- **`FullTextSynonymResolver`**: Similar to `FullTextResolver` yet searching the [`concept_synonym`](https://ohdsi.github.io/CommonDataModel/cdm54.html#concept_synonym) table
- **`PartialLabelResolver`**: Substring match. Is the search string `"Hodgkin lymphoma"` a partial component of any [`concept`](https://ohdsi.github.io/CommonDataModel/cdm54.html#concept)?
- **`PartialSynonymResolver`**: Similar to `PartialLabelResolver` yet searching the [`concept_synonym`](https://ohdsi.github.io/CommonDataModel/cdm54.html#concept_synonym) table
- **`FullTextResolver`**: Matches irrespective of word order. Not as relevant for the example above but relevant for others (e.g., "Kidney Cancer" -> "Cancer of Kidney").
- **`FullTextSynonymResolver`**: Similar to `FullTextResolver` yet searching the [`concept_synonym`](https://ohdsi.github.io/CommonDataModel/cdm54.html#concept_synonym) table
- **`EmbeddingResolver`**: Vector-similarity match, appended to the pipeline only when `omop-graph[emb]` (`omop-emb`) is installed.

!!! tip

Traversing each of the resolvers one by one can be an exhaustive search. The `ResolverPipeline` therefore offers a `stop_after_resolver` option. If set, retrieval from the DB stops after that resolver has concluded. The resolvers are ordered based on their confidence as above (i.e. **`ExactLabelResolver`** >> **`ExactSynonymResolver`** >> etc.)
Traversing each of the resolvers one by one can be an exhaustive search. The `ResolverPipeline` therefore offers a `stop_after_resolver` option. If set, retrieval from the DB stops after that resolver has concluded. `ALL_RESOLVERS`, the default sequence, is ordered exactly as listed above (i.e. **`ExactLabelResolver`** >> **`ExactSynonymResolver`** >> **`PartialLabelResolver`** >> ... >> **`EmbeddingResolver`**).

```python
from omop_alchemy.cdm.query import ConceptFilter
Expand Down
4 changes: 2 additions & 2 deletions docs/usage/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,5 +57,5 @@ omop-graph relationship-classification --pred-class-dir <PATH_TO_CSV_DIR>

| Option | Short | Type | Default | Description |
| :--- | :--- | :--- | :--- | :--- |
| **`--pred-class-dir`** | | `String` | **Required** | Path to the directory containing the classification CSVs. |
| **`--verbose`** | `-v` | `Count` | `0` | Increase logging verbosity (use `-v` or `-vv`). |
| **`--pred-class-dir`** | | `String` | `None` (bundled CSVs) | Path to a directory of classification CSVs, overriding the bundled defaults. |
| **`--verbose`**{: title="Global option, not specific to this subcommand — see the note above." } | `-v` | `Count` | `0` | Increase logging verbosity (use `-v` or `-vv`). Global option; must precede the subcommand name (see note above). |
18 changes: 17 additions & 1 deletion docs/usage/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,14 +20,30 @@ The full suite runs against an **in-memory SQLite mock CDM** (`tests/fixtures/mo

The grounding test suite is structured with parametrized cases so each clinical term is a separate pytest case for easier isolation and debugging.

### PostgreSQL-only integration suite

A separate suite, tagged with the `db_dialect` marker and excluded from the
default run (`pytest.toml`'s `-m "not db_dialect"`), runs against a real
PostgreSQL database via oa-configurator's test infrastructure. This is where
schema-drift protection and split-connection behavior are covered:
`test_schema_provenance_guard.py`, `test_vocab_split_connection.py`,
`test_oaklib_schema_awareness.py`, `test_fulltext_vocab_schema_postgres.py`,
`test_predicate_flags.py`, `test_relationship_classification.py`.

## Running Tests

Run all tests:
Run all tests (SQLite suite only, the default):

```bash
pytest
```

Include the PostgreSQL-only suite:

```bash
pytest -m db_dialect
```

Run one file:

```bash
Expand Down
3 changes: 3 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -80,6 +80,9 @@ omop-graph = "omop_graph.cli:app"
[project.entry-points."omop.config"]
omop_graph = "omop_graph.config:OmopGraphConfig"

[project.entry-points."oaklib.plugins"]
omop = "omop_graph.oaklib_interface.omop_implementation:OMOPAlchemyImplementation"

[build-system]
requires = ["hatchling", "hatch-vcs"]
build-backend = "hatchling.build"
Expand Down
2 changes: 1 addition & 1 deletion pytest.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[pytest]
testpaths = ["tests"]
addopts = ["-rf", "-rx", "--disable-pytest-warnings"]
addopts = ["-rf", "-rx", "--disable-pytest-warnings", "-m", "not db_dialect"]
log_cli = true
log_cli_level = "DEBUG"
log_cli_format = "%(asctime)s | %(name)s | %(levelname)s | %(message)s"
Expand Down
28 changes: 5 additions & 23 deletions scripts/benchmarks/benchmark.py
Original file line number Diff line number Diff line change
Expand Up @@ -13,9 +13,6 @@
from typing import Any, Dict, List, Optional, Sequence, Tuple, cast, Annotated
import typer

import sqlalchemy as sa
from sqlalchemy.orm import sessionmaker

import numpy as np
from oa_configurator import Resolver, ResolvedModel, ResolvedProvider
from omop_emb.config import (
Expand All @@ -25,7 +22,7 @@
parse_metric_type,
)
from omop_emb.interface import EmbeddingRole
from omop_emb.backends import EmbeddingBackend, resolve_backend_from_resolved_vector_store
from omop_emb.backends import EmbeddingBackend, open_vector_store_writer
from omop_emb.backends.index_config import index_config_from_index_type
from omop_graph.config import OmopGraphConfig
from omop_graph.extensions.emb import get_embedding_writer_interface, MissingExtensionError
Expand All @@ -45,7 +42,7 @@
PartialLabelResolver,
PartialSynonymResolver,
)
from omop_graph.db.session import make_engine
from omop_graph.db.session import resolve_cdm_database
app = typer.Typer()


Expand Down Expand Up @@ -163,29 +160,14 @@ def load_cases(path: Path) -> List[BenchmarkCase]:
raise TypeError(f"Unsupported benchmark case file shape: {type(payload).__name__}")


def build_session_factory() -> sessionmaker:
"""Build a SQLAlchemy session factory via oa-configurator."""
return sessionmaker(bind=make_engine(), future=True)


def build_engine() -> sa.Engine:
"""Build a SQLAlchemy engine via oa-configurator."""
return make_engine()


def build_knowledge_graph() -> KnowledgeGraph:
"""Create a KnowledgeGraph backed by the live OMOP CDM database."""
return KnowledgeGraph(cdm_engine=make_engine())


def build_embedding_knowledge_graph(
embedding_metric: MetricType,
resolved_model: ResolvedModel,
backend: EmbeddingBackend,
) -> KnowledgeGraph:
"""Create a KnowledgeGraph with embedding support configured."""

cdm_engine = make_engine()
cdm_engine = resolve_cdm_database().create_engine()
config = KnowledgeGraphEmbeddingConfiguration(
metric_type=embedding_metric,
backend=backend,
Expand Down Expand Up @@ -592,7 +574,7 @@ def run_benchmark(
"via `omop-config configure omop_graph`."
)
resolved_vector_store = Resolver.from_active_config().resolve_vector_store(resolved_name)
embedding_backend = resolve_backend_from_resolved_vector_store(resolved_vector_store)
embedding_backend = open_vector_store_writer(resolved_vector_store)

embedding_kg = build_embedding_knowledge_graph(
embedding_metric=resolved_embedding_metric_type,
Expand All @@ -616,7 +598,7 @@ def run_benchmark(
for case in cases
}

kg = build_knowledge_graph()
kg = KnowledgeGraph(cdm_engine=resolve_cdm_database().create_engine())
configs = build_grounded_configs()

errors: Dict[str, str] = {}
Expand Down
4 changes: 2 additions & 2 deletions scripts/benchmarks/enrich_gold_standard.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@

from sqlalchemy import text

from omop_graph.db.session import make_engine
from omop_graph.db.session import resolve_cdm_database

DOMAIN_QUERY = text(
"""
Expand Down Expand Up @@ -148,7 +148,7 @@ def main() -> None:
levels = [int(x) for x in args.levels.split(",")]

payload = json.loads(args.cases_file.read_text())
engine = make_engine()
engine = resolve_cdm_database().create_engine()

total_stats = Stats.empty_for(levels)
for bucket_name, cases in payload.items():
Expand Down
12 changes: 6 additions & 6 deletions scripts/benchmarks/trace_example.py
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@
)

from omop_graph.config import OmopGraphConfig
from omop_graph.db.session import make_engine
from omop_graph.db.session import resolve_cdm_database
from omop_graph.extensions.emb import get_embedding_writer_interface
from omop_graph.extensions.omop_alchemy import PredicateKind
from omop_alchemy.cdm.query import ConceptFilter
Expand Down Expand Up @@ -101,10 +101,10 @@ def _main(

def _build_kg(metric_type=None, embedding_model: Optional[str] = None) -> KnowledgeGraph:
"""Build a KG with embedding support, resolved from OmopGraphConfig."""
cdm_engine = make_engine()
cdm_engine = resolve_cdm_database().create_engine()
try:
from oa_configurator import Resolver
from omop_emb.backends import resolve_backend_from_resolved_vector_store
from omop_emb.backends import open_vector_store_writer
from omop_emb.config import MetricType

resolved_metric = metric_type if metric_type is not None else MetricType.COSINE
Expand All @@ -118,7 +118,7 @@ def _build_kg(metric_type=None, embedding_model: Optional[str] = None) -> Knowle
resolver = Resolver.from_active_config()
resolved_model = resolver.resolve_model(resolved_model_name)
resolved_vector_store = resolver.resolve_vector_store(cfg.vector_store_name)
backend = resolve_backend_from_resolved_vector_store(resolved_vector_store)
backend = open_vector_store_writer(resolved_vector_store)

emb_config = KnowledgeGraphEmbeddingConfiguration(
metric_type=resolved_metric,
Expand Down Expand Up @@ -1719,7 +1719,7 @@ def panel_svg(
else:
selected = cases

kg = KnowledgeGraph(cdm_engine=make_engine())
kg = KnowledgeGraph(cdm_engine=resolve_cdm_database().create_engine())
rel_class_cache: Dict[int, Dict] = {}

out_dir = trace_dir_pl / "plots" / "panel"
Expand Down Expand Up @@ -1782,7 +1782,7 @@ def graph_svg(
else:
selected = cases

kg = KnowledgeGraph(cdm_engine=make_engine())
kg = KnowledgeGraph(cdm_engine=resolve_cdm_database().create_engine())

out_dir = trace_dir_pl / "plots" / "graph"
out_dir.mkdir(parents=True, exist_ok=True)
Expand Down
Loading
Loading