Need benchmarks for the following: - Agent performance (including token spend) in TerminalBench and DeepSWE - TUI performance under heavy load (RAM usage, paint latency)
Need benchmarks for the following: