Evidence record
Prompt Schema Adaptation was rejected after 9/10 hard gates passed. Prompt Schema Adaptation · authenticated recorded trial
● Measured evidenceClassification measured
Procedure Generate a complete migration patch with the catalog-discovered NVIDIA model, execute it in a sibling Nebius Sandbox branch from the shared checkpoint, and classify all deterministic hard gates.
Observed result 9 of 10 hard gates passed; disposition rejected; measured branch duration 1554 ms (sample size 1).
Source revision b8f912021bc76a343b9bcf43bbd80f207ba253d3
Sandbox operation 01a07b2c-0d9c-7543-8ea6-c304951673ac
Artifact replay/live_20260907092102_streaming_retry_de3e3674/artifacts/sha256/37a68386809fbb17b3e8d9d00bd3ade14f645b1e7185760d123f4d17b2cd2bee
Content SHA-256 37a68386809fbb17b3e8d9d00bd3ade14f645b1e7185760d123f4d17b2cd2bee
Evidence IDs ev_2_build, ev_2_original_tests, ev_2_migration_tests, ev_2_schema, ev_2_tool_calls, ev_2_prompt_regression, ev_2_secret_scan, ev_2_security, ev_2_nebius_runtime, ev_2_evidence_completeness, ev_2_duration
Executed test logs and proposed diff Hash-verified recorded artifacts. Rejected patches remain visible for inspection, not approval to ship.
Executed test logs build: PASS · exit 0 · 1 ms
original-tests: PASS · exit 0 · 177 ms
----------------------------------------------------------------------
Ran 1 test in 0.000s
OK
schema: PASS · exit 0 · 170 ms
----------------------------------------------------------------------
Ran 2 tests in 0.000s
OK
tool-calls: PASS · exit 0 · 167 ms
----------------------------------------------------------------------
Ran 1 test in 0.000s
OK
streaming-retry: FAIL · exit 1 · 170 ms
======================================================================
FAIL: test_completion_usage_is_preserved (test_streaming_retry.StreamingAndRetry.test_completion_usage_is_preserved)
----------------------------------------------------------------------
Traceback (most recent call last):
File "/tmp/portverdict/test_streaming_retry.py", line 7, in test_completion_usage_is_preserved
self.assertEqual(adapter.normalize_event(event), {"choices":[],"usage":{"prompt_tokens":11,"completion_tokens":4,"total_tokens":15}})
AssertionError: {'cho[23 chars]rompt_tokens': 0, 'completion_tokens': 0, 'total_tokens': 0}} != {'cho[23 chars]rompt_tokens': 11, 'completion_tokens': 4, 'total_tokens': 15}}
{'choices': [],
- 'usage': {'completion_tokens': 0, 'prompt_tokens': 0, 'total_tokens': 0}}
? ^ ^ ^
+ 'usage': {'completion_tokens': 4, 'prompt_tokens': 11, 'total_tokens': 15}}
? ^ ^^ ^^
----------------------------------------------------------------------
Ran 2 tests in 0.001s
FAILED (failures=1)
security: PASS · exit 0 · 168 ms
----------------------------------------------------------------------
Ran 2 tests in 0.000s
OK
secret-scan: PASS · exit 0 · 0 ms
harness-integrity: PASS · exit 0 · 0 ms
Inspect candidate diff diff --git a/adapter.py b/adapter.py
--- a/adapter.py
+++ b/adapter.py
@@ -1,34 +1,61 @@
-import json
-
-def normalize_event(event):
- if event.get("type") == "chat.completion.chunk":
- return event
- if event.get("type") == "response.output_text.delta":
- return {"choices": [{"delta": {"content": event.get("delta", "")}}]}
- if event.get("type") == "response.output_item.added":
- call = normalize_tool_call(event.get("item", {}))
- return {"choices": [{"delta": {"tool_calls": [call]}}]} if call else {"choices": []}
- return {"choices": []}
-
-def parse_structured(text):
- raw = text.strip()
- if raw.startswith("```"):
- lines = raw.splitlines()
- if lines[0].strip() not in ("```", "```json") or lines[-1].strip() != "```":
- raise ValueError("invalid JSON fence")
- raw = "\n".join(lines[1:-1]).strip()
- data = json.loads(raw)
- if not isinstance(data, dict) or set(data) != {"answer"} or not isinstance(data["answer"], str) or not data["answer"].strip():
- raise ValueError("expected exactly one non-empty answer")
- return data
-
-def normalize_tool_call(item):
- if item.get("type") != "function_call":
- return None
- arguments = item.get("arguments")
- if not isinstance(arguments, str) or not isinstance(json.loads(arguments), dict):
- raise ValueError("function arguments must be a JSON object string")
- return {"id": item.get("id"), "type": "function", "function": {"name": item.get("name"), "arguments": arguments}}
-
-def retry_delay(status, attempt):
- return min(250 * (2 ** attempt), 2000)
+import json
+
+def normalize_event(event):
+ if not isinstance(event, dict):
+ return {"choices": []}
+ etype = event.get("type")
+ if etype == "chat.completion.chunk":
+ return event
+ if etype == "response.output_text.delta":
+ delta = event.get("delta", "")
+ return {"choices": [{"delta": {"content": delta}}]}
+ if etype == "response.output_item.added":
+ item = event.get("item", {})
+ if not isinstance(item, dict):
+ return {"choices": []}
+ call = normalize_tool_call(item)
+ if call is None:
+ return {"choices": []}
+ return {"choices": [{"delta": {"tool_calls": [call]}}]}
+ if etype == "response.completed":
+ usage = event.get("usage", {})
+ if not isinstance(usage, dict):
+ usage = {}
+ return {
+ "choices": [],
+ "usage": {
+ "prompt_tokens": usage.get("input_tokens", 0),
+ "completion_tokens": usage.get("output_tokens", 0),
+ "total_tokens": usage.get("total_tokens", 0)
+ }
+ }
+ return {"choices": []}
+
+def parse_structured(text):
+ raw = text.strip()
+ if raw.startswith("```"):
+ lines = raw.splitlines()
+ if lines[0].strip() not in ("```", "```json") or lines[-1].strip() != "```":
+ raise ValueError("invalid JSON fence")
+ raw = "\n".join(lines[1:-1]).strip()
+ data = json.loads(raw)
+ if not isinstance(data, dict) or set(data) != {"answer"} or not isinstance(data["answer"], str) or not data["answer"].strip():
+ raise ValueError("expected exactly one non-empty answer")
+ return data
+
+def normalize_tool_call(item):
+ if item.get("type") != "function_call":
+ return None
+ arguments = item.get("arguments")
+ if not isinstance(arguments, str) or not isinstance(json.loads(arguments), dict):
+ raise ValueError("function arguments must be a JSON object string")
+ return {"id": item.get("id"), "type": "function", "function": {"name": item.get("name"), "arguments": arguments}}
+
+def retry_delay(status, attempt):
+ if not isinstance(status, int) or not isinstance(attempt, int):
+ return None
+ if status == 429 or 500 <= status < 600:
+ delay = 250 * (2 ** attempt)
+ return delay if delay <= 2000 else 2000
+ return None
+
Raw log artifact Measured scope boundary This evidence supports one bounded recorded trial only. It is not a statistical claim and does not eliminate the need for patch review.
← Return to branch workflow