Evidence record
Minimal Compatibility was rejected after 9/10 hard gates passed. Minimal Compatibility · authenticated recorded trial
● Measured evidenceClassification measured
Procedure Generate a complete migration patch with the catalog-discovered NVIDIA model, execute it in a sibling Nebius Sandbox branch from the shared checkpoint, and classify all deterministic hard gates.
Observed result 9 of 10 hard gates passed; disposition rejected; measured branch duration 1520 ms (sample size 1).
Source revision b8f912021bc76a343b9bcf43bbd80f207ba253d3
Sandbox operation 01a07b2c-0d94-7275-bc61-20369fe4c945
Artifact replay/live_20260907092102_streaming_retry_de3e3674/artifacts/sha256/d19f462f03d14ab82495e5a0b43272c35a040e5a82b64893bbf829434ae4104a
Content SHA-256 d19f462f03d14ab82495e5a0b43272c35a040e5a82b64893bbf829434ae4104a
Evidence IDs ev_1_build, ev_1_original_tests, ev_1_migration_tests, ev_1_schema, ev_1_tool_calls, ev_1_prompt_regression, ev_1_secret_scan, ev_1_security, ev_1_nebius_runtime, ev_1_evidence_completeness, ev_1_duration
Executed test logs and proposed diff Hash-verified recorded artifacts. Rejected patches remain visible for inspection, not approval to ship.
Executed test logs build: PASS · exit 0 · 1 ms
original-tests: PASS · exit 0 · 180 ms
----------------------------------------------------------------------
Ran 1 test in 0.000s
OK
schema: PASS · exit 0 · 173 ms
----------------------------------------------------------------------
Ran 2 tests in 0.000s
OK
tool-calls: PASS · exit 0 · 171 ms
----------------------------------------------------------------------
Ran 1 test in 0.000s
OK
streaming-retry: FAIL · exit 1 · 181 ms
======================================================================
FAIL: test_completion_usage_is_preserved (test_streaming_retry.StreamingAndRetry.test_completion_usage_is_preserved)
----------------------------------------------------------------------
Traceback (most recent call last):
File "/tmp/portverdict/test_streaming_retry.py", line 7, in test_completion_usage_is_preserved
self.assertEqual(adapter.normalize_event(event), {"choices":[],"usage":{"prompt_tokens":11,"completion_tokens":4,"total_tokens":15}})
AssertionError: {'cho[23 chars]rompt_tokens': 0, 'completion_tokens': 0, 'total_tokens': 0}} != {'cho[23 chars]rompt_tokens': 11, 'completion_tokens': 4, 'total_tokens': 15}}
{'choices': [],
- 'usage': {'completion_tokens': 0, 'prompt_tokens': 0, 'total_tokens': 0}}
? ^ ^ ^
+ 'usage': {'completion_tokens': 4, 'prompt_tokens': 11, 'total_tokens': 15}}
? ^ ^^ ^^
----------------------------------------------------------------------
Ran 2 tests in 0.001s
FAILED (failures=1)
security: PASS · exit 0 · 164 ms
----------------------------------------------------------------------
Ran 2 tests in 0.000s
OK
secret-scan: PASS · exit 0 · 0 ms
harness-integrity: PASS · exit 0 · 0 ms
Inspect candidate diff diff --git a/adapter.py b/adapter.py
--- a/adapter.py
+++ b/adapter.py
@@ -1,34 +1,49 @@
-import json
-
-def normalize_event(event):
- if event.get("type") == "chat.completion.chunk":
- return event
- if event.get("type") == "response.output_text.delta":
- return {"choices": [{"delta": {"content": event.get("delta", "")}}]}
- if event.get("type") == "response.output_item.added":
- call = normalize_tool_call(event.get("item", {}))
- return {"choices": [{"delta": {"tool_calls": [call]}}]} if call else {"choices": []}
- return {"choices": []}
-
-def parse_structured(text):
- raw = text.strip()
- if raw.startswith("```"):
- lines = raw.splitlines()
- if lines[0].strip() not in ("```", "```json") or lines[-1].strip() != "```":
- raise ValueError("invalid JSON fence")
- raw = "\n".join(lines[1:-1]).strip()
- data = json.loads(raw)
- if not isinstance(data, dict) or set(data) != {"answer"} or not isinstance(data["answer"], str) or not data["answer"].strip():
- raise ValueError("expected exactly one non-empty answer")
- return data
-
-def normalize_tool_call(item):
- if item.get("type") != "function_call":
- return None
- arguments = item.get("arguments")
- if not isinstance(arguments, str) or not isinstance(json.loads(arguments), dict):
- raise ValueError("function arguments must be a JSON object string")
- return {"id": item.get("id"), "type": "function", "function": {"name": item.get("name"), "arguments": arguments}}
-
-def retry_delay(status, attempt):
- return min(250 * (2 ** attempt), 2000)
+import json
+
+def normalize_event(event):
+ if event.get("type") == "chat.completion.chunk":
+ return event
+ if event.get("type") == "response.output_text.delta":
+ return {"choices": [{"delta": {"content": event.get("delta", "")}}]}
+ if event.get("type") == "response.output_item.added":
+ call = normalize_tool_call(event.get("item", {}))
+ if call:
+ return {"choices": [{"delta": {"tool_calls": [call]}}]}
+ return {"choices": []}
+ if event.get("type") == "response.completed":
+ usage = event.get("usage", {})
+ return {
+ "choices": [],
+ "usage": {
+ "prompt_tokens": usage.get("input_tokens", 0),
+ "completion_tokens": usage.get("output_tokens", 0),
+ "total_tokens": usage.get("total_tokens", 0)
+ }
+ }
+ return {"choices": []}
+
+def parse_structured(text):
+ raw = text.strip()
+ if raw.startswith("```"):
+ lines = raw.splitlines()
+ if lines[0].strip() not in ("```", "```json") or lines[-1].strip() != "```":
+ raise ValueError("invalid JSON fence")
+ raw = "\n".join(lines[1:-1]).strip()
+ data = json.loads(raw)
+ if not isinstance(data, dict) or set(data) != {"answer"} or not isinstance(data["answer"], str) or not data["answer"].strip():
+ raise ValueError("expected exactly one non-empty answer")
+ return data
+
+def normalize_tool_call(item):
+ if item.get("type") != "function_call":
+ return None
+ arguments = item.get("arguments")
+ if not isinstance(arguments, str) or not isinstance(json.loads(arguments), dict):
+ raise ValueError("function arguments must be a JSON object string")
+ return {"id": item.get("id"), "type": "function", "function": {"name": item.get("name"), "arguments": arguments}}
+
+def retry_delay(status, attempt):
+ if status == 429 or 500 <= status < 600:
+ return min(250 * (2 ** attempt), 2000)
+ return None
+
Raw log artifact Measured scope boundary This evidence supports one bounded recorded trial only. It is not a statistical claim and does not eliminate the need for patch review.
← Return to branch workflow