Repository navigation
Enrich LLM review with compiler-linked cross-package evidence - #8
Merged
Merged
Conversation
I3eg1nner
marked this pull request as ready for review
September 26, 2026 08:00
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem and resulting behavior
The LLM client stored dataflow in its bundle but mainly sent nearby source windows. A cross-package helper's behavior was missing even when the analyzer had already resolved the call.
Completed IR functions/callbacks now carry AST source ranges. The review helper matches each finding to its exact reported route/path, follows IR calls, and sends complete related source spans, unit identities, operations, raw bindings and declared models. Per-finding attachment IDs support exact cross-file citations. The whole project input set/hash is rechecked before and after a request, including configuration and dependencies. Missing identities, changed inputs, unsupported ranges and budgets remain explicit gaps.
Validation
.envrequests on the same three findings: windows produced 3 needs_review opinions; enriched evidence produced 3 supported_concern opinions with 5 valid cross-file quotes. Independent review confirmed all 8 enriched quotes and the helper/parameter explanations against the actual IR.experiments/llm_evidence/: prompt tokens 3,040 → 13,219; total tokens 6,805 → 17,807. Credentials, endpoint and model identity are excluded.Boundaries
This is evidence enrichment for existing findings, not autonomous browsing or new vulnerability discovery. The helper checks trusted-report consistency and source freshness; it does not recompute official bindings. All model opinions remain llm_unverified. Three synthetic positive findings and one call per variant do not establish overall precision, recall or lower false-positive rates. Ordinary syntax hints without compiler-linked evidence retain windows.
Final independent delivery acceptance
Accepted head:
1d7c41e7158bb80c2805c6fa056d9b232426ad27.All 9 checks passed: native artifacts, build and compatibility.
An independent subagent downloaded all three artifacts, verified archive/binary/build-info identities and all five acceptance suites. Each platform passes LLM context 4, compilation contexts 7, cross-package 9, cmark 14, extracted API helper 16 and unchanged Luna runtime 15; semantic cases are Linux 16, macOS 13, Windows 14. Actual helper/source selection is 3/2/3 functions and attachments for the three routes.
Linux/macOS helper raw SHA matches the reviewed file. Windows raw SHA matches its own acceptance report; exact CRLF-to-LF normalization matches the reviewed bytes, with no other difference. Both identities are recorded below. Linux/macOS observe 256-query budget exhaustion; Windows instead reaches its 60-second worker ceiling and conservatively returns incomplete. No Windows 256-query measurement is claimed.
Final artifact identity ledger
{ "head_sha": "1d7c41e7158bb80c2805c6fa056d9b232426ad27", "run_id": 36227446571, "platforms": [ { "platform": "linux-x86_64", "archive": "moon-audit-0.5.0-dev-linux-x86_64.zip", "archive_sha256": "ef72854bbd8fccf208e5d9a63db096eb820ce33ef4578b1e4fdf7e1033bcca33", "binary_sha256": "f327995494630a9ec4a5933f869878a2b5908ced6bbb5238fe0f1640586c733e", "acceptance": { "cross-package": 9, "semantic": 16, "cmark": 14, "compilation-context": 7, "llm-context": 4 }, "reached_body_limit": { "kind": "binding_queries", "query_limit_observed": true, "binding_queries": 256 }, "budget_case_seconds": 33.904, "real_package_runtime_tests": 15, "unsupported_body_query_count": 1, "unsupported_body_seconds": 0.891, "context_run_seconds": { "inline_test_only": 1.326, "production_with_all_test_roles": 1.658, "cfg_false": 0.637, "cfg_js": 0.647, "reached_conditional_helper": 1.689, "unsupported_body_260_calls": 0.891, "shared_helper_resolution_failure": 4.059 }, "review_script_sha256": "9ee892542a94137533cc557d0598f58e66a822796bd4633501621d99719ce6c8", "review_script_lf_normalized_sha256": "9ee892542a94137533cc557d0598f58e66a822796bd4633501621d99719ce6c8", "llm_context_summary": { "/raw": { "functions": 3, "attachments": 3, "status": "ready" }, "/fake": { "functions": 2, "attachments": 2, "status": "ready" }, "/second-value": { "functions": 3, "attachments": 3, "status": "ready" } }, "api_helper_tests": { "source": 16, "extracted": 16, "log_sha256": "86d16a0d676146235694fd8a5cc912a33b6c2486c5d0385940b0bf31d50fc910", "summaries": [ "2026-09-26T07:39:57.2341235Z Ran 16 tests in 15.998s", "2026-09-26T07:44:42.2352765Z Ran 16 tests in 16.022s" ] }, "artifact_id": 10901024368, "job_id": 108364070497 }, { "platform": "macos-arm64", "archive": "moon-audit-0.5.0-dev-macos-arm64.zip", "archive_sha256": "d722c7849ea1002ea59c5501d74ebb72b1d4b576d6f0692d97997d02465bb65c", "binary_sha256": "97c2dffee6fcc1127ce1f4cef5fca19b0db88aaf10650e83cf26eef54a85c826", "acceptance": { "cross-package": 9, "semantic": 13, "cmark": 14, "compilation-context": 7, "llm-context": 4 }, "reached_body_limit": { "kind": "binding_queries", "query_limit_observed": true, "binding_queries": 256 }, "budget_case_seconds": 47.916, "real_package_runtime_tests": 15, "unsupported_body_query_count": 1, "unsupported_body_seconds": 1.238, "context_run_seconds": { "inline_test_only": 2.748, "production_with_all_test_roles": 2.348, "cfg_false": 1.156, "cfg_js": 1.028, "reached_conditional_helper": 2.213, "unsupported_body_260_calls": 1.238, "shared_helper_resolution_failure": 6.216 }, "review_script_sha256": "9ee892542a94137533cc557d0598f58e66a822796bd4633501621d99719ce6c8", "review_script_lf_normalized_sha256": "9ee892542a94137533cc557d0598f58e66a822796bd4633501621d99719ce6c8", "llm_context_summary": { "/raw": { "functions": 3, "attachments": 3, "status": "ready" }, "/fake": { "functions": 2, "attachments": 2, "status": "ready" }, "/second-value": { "functions": 3, "attachments": 3, "status": "ready" } }, "api_helper_tests": { "source": 16, "extracted": 16, "log_sha256": "4e649ebda1538bdbed066ce9a25fcf86654663ee55892ba5d9324b3ddd19c085", "summaries": [ "2026-09-26T07:39:59.4919090Z Ran 16 tests in 15.664s", "2026-09-26T07:44:37.0215470Z Ran 16 tests in 16.716s" ] }, "artifact_id": 10901218991, "job_id": 108364070493, "macos_memory_acceptance": "passed" }, { "platform": "windows-x86_64", "archive": "moon-audit-0.5.0-dev-windows-x86_64.zip", "archive_sha256": "85220ffd080a29d4c0a55e33c07f6f3c15b628af4ce9111758e88bcbece28672", "binary_sha256": "1f240243fbf42cec0638279791fbcd6982c1ad8602dec16d3211caf584da8de5", "acceptance": { "cross-package": 9, "semantic": 14, "cmark": 14, "compilation-context": 7, "llm-context": 4 }, "reached_body_limit": { "kind": "worker_timeout", "query_limit_observed": false }, "budget_case_seconds": 61.187, "real_package_runtime_tests": 15, "unsupported_body_query_count": 1, "unsupported_body_seconds": 2.594, "context_run_seconds": { "inline_test_only": 3.094, "production_with_all_test_roles": 4.531, "cfg_false": 1.86, "cfg_js": 1.891, "reached_conditional_helper": 4.734, "unsupported_body_260_calls": 2.594, "shared_helper_resolution_failure": 10.797 }, "review_script_sha256": "7d3ede302be93efce1615716ff855f143028c831914308288552fb93b7280f3c", "review_script_lf_normalized_sha256": "9ee892542a94137533cc557d0598f58e66a822796bd4633501621d99719ce6c8", "llm_context_summary": { "/raw": { "functions": 3, "attachments": 3, "status": "ready" }, "/fake": { "functions": 2, "attachments": 2, "status": "ready" }, "/second-value": { "functions": 3, "attachments": 3, "status": "ready" } }, "api_helper_tests": { "source": 16, "extracted": 16, "log_sha256": "599afd9a8fdd40beef439584ed2dd9a2ac98e44dee68e49a602239bad8169655", "summaries": [ "2026-09-26T07:40:08.5973851Z Ran 16 tests in 19.810s", "2026-09-26T07:46:28.1598171Z Ran 16 tests in 19.099s" ] }, "artifact_id": 10901139597, "job_id": 108364070397 } ], "status": "passed", "blocking_findings": [], "notes": [ "Windows helper uses CRLF checkout line endings; its raw hash equals its own acceptance record. Normalizing only CRLF to LF reproduces the exact reviewed helper bytes and SHA on Linux/macOS.", "Windows reachable-body stress case ended at the 60-second semantic worker timeout; it did not establish the 256-query limit. Exit 2, errors, compiler verification and syntax hint were preserved with no semantic finding." ] }