Bidlo Status Frontend

Chat tool telemetry

Find slow, failing, and repetitive agent tools without exposing prompts, arguments, or outputs.

Detailed production timing active

6101 of 6166 saved tool calls include execution duration.

98.9% timing coverage

Tool calls

5

-17% vs previous 7d4 saved production turns

p50 latency

490 ms

+17% vs previous 7d5 timed calls

p95 latency

599 ms

-6% vs previous 7d5 timed calls

Tool failures

20.0%

+3.3 pts vs previous 7d1 excluding NoResults

No result

0.0%

No change vs previous 7d0 product outcomes

Calls / turn

1.3

+4% vs previous 7d75.0% of turns at 5+ calls

What changed

Material regressions ranked by severity and affected call volume.

Current 7 days compared with the previous equal window

Reliability regression

create_database_view failures increased

20.0% now vs 16.7% previously across 5 calls; InvalidArguments is the leading code.

Inspect recent failures, then check the tool handler and upstream dependency logs.

Investigate

Agent loop pressure

More turns are using five or more tools

75.0% now vs 60.0% previously (3 turns).

Inspect the affected turn sequences for repeated or avoidable tool selection.

Investigate

Regression checks require at least 3 calls in both windows. Latency is flagged at +20%; failure and no-result rates are flagged at +1 and +3 points respectively.

Execution volume

Successful calls, no-results, and failures combined.

Calls
2 calls on 2026-10-043 calls on 2026-10-06
Oct 1Oct 4Oct 7

Payload pressure

How often results exceed the 20 KB inline cap.

5.0%

handle rows / all calls

309 stored handles in this period. This is a payload-size proxy, not a one-to-one call classification.

Tool comparison

Each change is measured against the previous equal window; latency appears only where detailed timing exists.

Sorted by current p95 latency
ToolCallsp50p95FailuresNo result
create_database_view5-17% vs prior490 ms+17% vs prior599 ms-6% vs prior20.0%+3.3 pts vs prior0.0%No change vs prior

Failure codes

Operational failures only; NoResults is excluded.

InvalidArguments1

Interpretation notes

The data has known edges that change what a chart means.

Blocked loops are invisible
Rate-limited and identical-argument calls are rejected before telemetry fires.
NoResults is not downtime
It remains a failed tool outcome, but is separated from runtime reliability here.
Preview volume lives elsewhere
Preview and threadless chats reach PostHog but are intentionally absent from Postgres.
Latency is partially instrumented
Saved tool parts provide production outcomes and volume; duration appears only on calls with detailed telemetry.
Partial turns are retained
They remain in volume counts but are excluded from failures and latency percentiles.
Comparisons use equal windows
Every delta compares the selected period with the immediately preceding period under the same tool and team filters.

Recent sanitized calls

Arguments, outputs, user IDs, and message content are never rendered.

Privacy safe
create_database_viewHJC490 msOct 6, 7:14 PM UTC

Outcome

Success

Thread

3d7ad6cb…

Tool call

chatcmpl…

Turn sequence

8 of 9

Turn state

Completed

create_database_viewHJC554 msOct 6, 7:02 PM UTC

Outcome

Success

Thread

033cbc16…

Tool call

chatcmpl…

Turn sequence

5 of 5

Turn state

Completed

create_database_viewHJC430 msOct 6, 6:56 PM UTC

Outcome

Success

Thread

50ac87df…

Tool call

call_io2…

Turn sequence

1 of 1

Turn state

Completed

create_database_viewHJC211 msOct 4, 4:23 PM UTC

Outcome

InvalidArguments

Thread

e1179f50…

Tool call

chatcmpl…

Turn sequence

10 of 12

Turn state

Completed

create_database_viewHJC599 msOct 4, 4:23 PM UTC

Outcome

Success

Thread

e1179f50…

Tool call

chatcmpl…

Turn sequence

11 of 12

Turn state

Completed

Source: Production chat_messages.tool_results (saved tool parts and detailed telemetry) Tool execution time starts after approval gates.