fix(ai): give the post-exit stderr drain room to be scheduled

The reaped-turn drain wait assumed the pipe was at EOF so the task
"returns at once", but the wait is really for the drain TASK to be
scheduled on a saturated runtime. Under the concurrent stress test (and
the orchestrator's parallel turns) that scheduling latency outran the
two-second bound, the tail read back empty, and a child that explained
itself on stderr surfaced as "(no output captured)". Raise the grace to
thirty seconds: it only has to outlast scheduler starvation while still
capping a genuinely wedged reader. Stress test green 5/5 locally.
This commit is contained in:
Kayshen-X 2026-08-10 06:10:38 +08:00
parent 70fe62e30c
commit c658706d04

View file

@ -116,9 +116,15 @@ const STDOUT_TAIL_CAP: usize = 8 * 1024;
const STDOUT_TAIL_LINES: usize = 256;
/// How long a reaped turn waits for the stderr drain to reach EOF
/// before formatting its failure message. Normally one scheduler round;
/// bounded so a wedged reader cannot hang a turn.
const STDERR_DRAIN_GRACE: Duration = Duration::from_secs(2);
/// before formatting its failure message. The child is already reaped so
/// its pipe is at EOF and the drain finishes the instant it is polled —
/// the wait is really for the drain TASK to be scheduled, not for I/O.
/// Under a saturated runtime (the orchestrator's parallel turns, or the
/// concurrent stress test) that scheduling latency occasionally ran past
/// a two-second bound and the tail read back empty, so this is generous:
/// it only has to outlast scheduler starvation, while still capping a
/// genuinely wedged reader (a grandchild holding the pipe open).
const STDERR_DRAIN_GRACE: Duration = Duration::from_secs(30);
/// `ChatProvider` impl that bridges to a CLI binary via stdio.
/// Construct via [`SubprocessProvider::for_cli`] or