Skip to content

Parse hcl in a reusable worker process - #448

Open
benlangfeld wants to merge 1 commit into
dflook:mainfrom
benlangfeld:hcl-parser-worker
Open

benlangfeld wants to merge 1 commit into
dflook:mainfrom
benlangfeld:hcl-parser-worker

Conversation

@benlangfeld

Copy link
Copy Markdown

Fixes #447.

terraform.hcl.load() started a fresh Python interpreter for every file it was asked to parse, to find out whether the file could be parsed, and then parsed the file again in the calling process. The interpreter start and the import hcl2/import lark it pulls in cost ~113ms, which load_module() pays once per .tf file in the root module.

The subprocess was added in f87a90c to mitigate hangs from unterminated strings, and that requirement still stands — a hang inside a re match can't be interrupted by an in-process timeout, so the parse does have to happen somewhere killable. It doesn't need a new interpreter per file though.

This runs the parse in a worker process that is reused across calls, and is only replaced when one actually has to be killed. A module pays for one worker instead of one per file.

Effect

On a module with 245 root .tf files (14k lines, 1.1MB), python-hcl2 7.3.1:

before after
load_module() 41.2s 0.6s

Output is identical (860 top-level blocks both ways). load_module() is called from terraform-backend and github_pr_comment, so on that module this is worth ~90s per terraform-plan run and ~170s per terraform-apply.

Notes

  • Parse failures are returned from the worker rather than raised, because lark's exceptions aren't picklable — raising them across the process boundary replaced the real parse error with cannot pickle 'module' object. The debug log keeps the same content it has today.
  • load()'s contract is unchanged: a timeout raises ValueError (so load_module() skips the file, as now), anything else falls back to fallback_parser. loads() still raises ValueError either way.
  • is_loadable() and the __main__ block existed only to support the subprocess and are gone. Nothing else referenced them.
  • fork is preferred for the worker so it inherits the already-imported parser; without it the worker starts an interpreter, but still only once.
  • Added tests/test_hcl.py covering the timeout kill and worker reuse. pytest tests passes (213 tests, excluding the ones needing network access to releases.hashicorp.com), as do ruff --config=.config/ruff.toml check and mypy.
  • I've left CHANGELOG.md alone — happy to add an entry in whatever form you'd like.

🤖 Generated with Claude Code

load() started a fresh Python interpreter for every file it was asked to
parse, so that a parse could be killed if it hung, and then parsed the
file again in the calling process. The interpreter start and the lark
import it pulls in cost about 113ms, which load_module() pays once per
file in the root module.

On a 245 file module that is 33s of process spawning to do 0.6s of
parsing, and load_module() is called twice per plan and four times per
apply, so it is around 90s of a plan and 170s of an apply.

Run the parse in a worker process that is reused across calls instead.
Killing a hung parse still works, because the worker is a separate
process that can be terminated, but a module now pays for one worker
rather than one per file: load_module() over those 245 files goes from
41.2s to 0.6s, with identical output.

The worker returns the reason a parse failed rather than raising, since
lark's exceptions can't be pickled back to the caller and the reason is
wanted for the debug log.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@benlangfeld
benlangfeld marked this pull request as ready for review September 24, 2026 16:38

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

HCL parsing spawns a Python interpreter per .tf file, costing ~90s per plan on a large module

1 participant