fix(ai): give the post-exit stderr drain room to be scheduled
The reaped-turn drain wait assumed the pipe was at EOF so the task "returns at once", but the wait is really for the drain TASK to be scheduled on a saturated runtime. Under the concurrent stress test (and the orchestrator's parallel turns) that scheduling latency outran the two-second bound, the tail read back empty, and a child that explained itself on stderr surfaced as "(no output captured)". Raise the grace to thirty seconds: it only has to outlast scheduler starvation while still capping a genuinely wedged reader. Stress test green 5/5 locally.
This commit is contained in:
parent
70fe62e30c
commit
c658706d04
|
|
@ -116,9 +116,15 @@ const STDOUT_TAIL_CAP: usize = 8 * 1024;
|
|||
const STDOUT_TAIL_LINES: usize = 256;
|
||||
|
||||
/// How long a reaped turn waits for the stderr drain to reach EOF
|
||||
/// before formatting its failure message. Normally one scheduler round;
|
||||
/// bounded so a wedged reader cannot hang a turn.
|
||||
const STDERR_DRAIN_GRACE: Duration = Duration::from_secs(2);
|
||||
/// before formatting its failure message. The child is already reaped so
|
||||
/// its pipe is at EOF and the drain finishes the instant it is polled —
|
||||
/// the wait is really for the drain TASK to be scheduled, not for I/O.
|
||||
/// Under a saturated runtime (the orchestrator's parallel turns, or the
|
||||
/// concurrent stress test) that scheduling latency occasionally ran past
|
||||
/// a two-second bound and the tail read back empty, so this is generous:
|
||||
/// it only has to outlast scheduler starvation, while still capping a
|
||||
/// genuinely wedged reader (a grandchild holding the pipe open).
|
||||
const STDERR_DRAIN_GRACE: Duration = Duration::from_secs(30);
|
||||
|
||||
/// `ChatProvider` impl that bridges to a CLI binary via stdio.
|
||||
/// Construct via [`SubprocessProvider::for_cli`] or
|
||||
|
|
|
|||
Loading…
Reference in a new issue