One session. Two phases.
Started at 09:04 UTC. An outer launcher timeout interrupted the process after 30 minutes. The same session resumed at 09:41 UTC and completed at 17:19 UTC. The model and original goal remained unchanged.
Exact benchmark prompt
Activity across the run
Click a 15-minute interval to jump into it. Bar height = recorded events, including provider requests and context updates.
Accounting, without double counting
Tool breakdown and coverage
Full event timeline
What this archive does—and does not—claim
- This is the tested Claude Code session, not the private Codex/OpenClaw orchestration conversation. The tested model is DeepSeek, not Claude.
- Tool invocations are deduplicated by tool-use ID across the main transcript and two subagent transcripts. A tool call can launch several shell commands; these are not counted as separate tool calls.
- Token and dollar totals use the last provider usage record per request ID. Initial zero-usage response headers and repeated assistant content blocks are not summed. Cached-input tokens are reported separately; reasoning tokens are a subset of output tokens, not additional tokens.
- API requests also include internal harness work (e.g. compaction), so request count differs from visible model-message count. Failed requests without usage may have unrecorded provider charges. Recorded cost is not an invoice.
- Tool duration is timestamp difference between invocation and result, not CPU execution time. Background tools may return before their process finishes.
- Images embedded in tool results are deduplicated by original-byte SHA-256, converted to WebP and resized to at most 1600×1200. They are historical session images, not later presentation-camera renders. Repeated views remain linked to each event.
- Runtime pointer bookkeeping (mode / last-prompt / atis-latch / queue-operation) is counted in source inventories but omitted from the interaction timeline. All assistant, user, system and attachment interaction records are retained. No claim is made that every filesystem screenshot was opened by the model.
- Public data are sanitized: credential-like values, authorization strings, environment fields and private identifiers are removed. Original benchmark logs remain unchanged. Exported JSON is a derived archive, not an unmodified raw-log leak.
- Use the final result record for the model’s own claims and acknowledged limitations; successful process exit is not an independent realism score.