No stage activity is simulated while evidence is being read.
Evidence record
Prompt Schema Adaptation was rejected after 8/10 hard gates passed.
Prompt Schema Adaptation · authenticated recorded trial
●Measured evidence
Classificationmeasured
Procedure
Generate a complete migration patch with the catalog-discovered NVIDIA model, execute it in a sibling Nebius Sandbox branch from the shared checkpoint, and classify all deterministic hard gates.
Observed result
8 of 10 hard gates passed; disposition rejected; measured branch duration 1555 ms (sample size 1).
Hash-verified recorded artifacts. Rejected patches remain visible for inspection, not approval to ship.
Executed test logs
build: PASS · exit 0 · 2 ms
original-tests: PASS · exit 0 · 184 ms
----------------------------------------------------------------------
Ran 1 test in 0.000s
OK
schema: PASS · exit 0 · 172 ms
----------------------------------------------------------------------
Ran 2 tests in 0.000s
OK
tool-calls: PASS · exit 0 · 169 ms
----------------------------------------------------------------------
Ran 1 test in 0.000s
OK
streaming-retry: PASS · exit 0 · 165 ms
----------------------------------------------------------------------
Ran 2 tests in 0.000s
OK
security: FAIL · exit 1 · 181 ms
======================================================================
FAIL: test_invalid_arguments_are_rejected (test_hidden.HiddenFalsifier.test_invalid_arguments_are_rejected)
----------------------------------------------------------------------
Traceback (most recent call last):
File "/tmp/portverdict/test_hidden.py", line 8, in test_invalid_arguments_are_rejected
with self.assertRaises((ValueError, TypeError)):
AssertionError: (<class 'ValueError'>, <class 'TypeError'>) not raised
----------------------------------------------------------------------
Ran 2 tests in 0.000s
FAILED (failures=1)
secret-scan: PASS · exit 0 · 0 ms
harness-integrity: PASS · exit 0 · 0 ms
Inspect candidate diff
diff --git a/adapter.py b/adapter.py
--- a/adapter.py
+++ b/adapter.py
@@ -1,34 +1,58 @@
-import json
-
-def normalize_event(event):
- if event.get("type") == "chat.completion.chunk":
- return event
- if event.get("type") == "response.output_text.delta":
- return {"choices": [{"delta": {"content": event.get("delta", "")}}]}
- if event.get("type") == "response.output_item.added":
- call = normalize_tool_call(event.get("item", {}))
- return {"choices": [{"delta": {"tool_calls": [call]}}]} if call else {"choices": []}
- if event.get("type") == "response.completed":
- usage = event.get("response", {}).get("usage", {})
- return {"choices": [], "usage": {"prompt_tokens": usage.get("input_tokens", 0), "completion_tokens": usage.get("output_tokens", 0), "total_tokens": usage.get("total_tokens", 0)}}
- return {"choices": []}
-
-def parse_structured(text):
- raw = text.strip()
- if raw.startswith("```"):
- lines = raw.splitlines()
- if lines[0].strip() not in ("```", "```json") or lines[-1].strip() != "```":
- raise ValueError("invalid JSON fence")
- raw = "\n".join(lines[1:-1]).strip()
- data = json.loads(raw)
- if not isinstance(data, dict) or set(data) != {"answer"} or not isinstance(data["answer"], str) or not data["answer"].strip():
- raise ValueError("expected exactly one non-empty answer")
- return data
-
-def normalize_tool_call(item):
- return None
-
-def retry_delay(status, attempt):
- if status != 429 and not 500 <= status <= 599:
- return None
- return min(250 * (2 ** attempt), 2000)
+import json
+
+def normalize_event(event):
+ if event.get("type") == "chat.completion.chunk":
+ return event
+ if event.get("type") == "response.output_text.delta":
+ return {"choices": [{"delta": {"content": event.get("delta", "")}}]}
+ if event.get("type") == "response.output_item.added":
+ call = normalize_tool_call(event.get("item", {}))
+ return {"choices": [{"delta": {"tool_calls": [call]}}]} if call else {"choices": []}
+ if event.get("type") == "response.completed":
+ usage = event.get("response", {}).get("usage", {})
+ return {"choices": [], "usage": {"prompt_tokens": usage.get("input_tokens", 0), "completion_tokens": usage.get("output_tokens", 0), "total_tokens": usage.get("total_tokens", 0)}}
+ return {"choices": []}
+
+def parse_structured(text):
+ raw = text.strip()
+ if raw.startswith("```"):
+ lines = raw.splitlines()
+ if lines[0].strip() not in ("```", "```json") or lines[-1].strip() != "```":
+ raise ValueError("invalid JSON fence")
+ raw = "\n".join(lines[1:-1]).strip()
+ data = json.loads(raw)
+ if not isinstance(data, dict) or set(data) != {"answer"} or not isinstance(data["answer"], str) or not data["answer"].strip():
+ raise ValueError("expected exactly one non-empty answer")
+ return data
+
+def normalize_tool_call(item):
+ if not isinstance(item, dict):
+ return None
+ t = item.get('type')
+ if t != 'function_call':
+ return None
+ i = item.get('id')
+ n = item.get('name')
+ a = item.get('arguments')
+ if not isinstance(i, str) or not i:
+ return None
+ if not isinstance(n, str) or not n:
+ return None
+ if isinstance(a, dict):
+ try:
+ a = json.dumps(a)
+ except Exception:
+ return None
+ elif isinstance(a, str):
+ try:
+ json.loads(a)
+ except Exception:
+ return None
+ else:
+ return None
+ return {'id': i, 'type': 'function', 'function': {'name': n, 'arguments': a}}
+
+def retry_delay(status, attempt):
+ if status != 429 and not 500 <= status <= 599:
+ return None
+ return min(250 * (2 ** attempt), 2000)