op-smoke's QueryEngine path can't send MiniMax's `thinking:{type:
"disabled"}` field, so reasoning models (M3 etc.) thought themselves out
of the token budget and couldn't be benchmarked end-to-end. Add a
non-streaming DirectClient (OPENPENCIL_SMOKE_DIRECT=1) that posts
openai-compat directly and disables thinking at the wire for MiniMax —
or any endpoint via OPENPENCIL_SMOKE_DISABLE_THINKING=1 (Volcengine
honors the same param) — so the production thinking-disable path can be
exercised headless without a GUI.