Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ jobs:
with:
python-version: ${{ matrix.python-version }}

- run: pip install -e ".[dev,http,apprise]"
- run: pip install -e ".[dev,http,apprise,otel]"

- run: ruff check .

Expand Down
26 changes: 26 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,32 @@ All notable changes to this project are documented in this file.

## Unreleased

- LLM calls and OpenTelemetry export:
- `logger.llm_call(model=, tokens_in=, tokens_out=, cost_usd=, latency_ms=,
finish_reason=)` records an LLM call with its numbers in the record's
first-class `llm` block. A value that breaks the contract is dropped with a
warning instead of raising.
- `OTLPTransport` (`pip install logquill[otel]`) exports agent tracing as real
OpenTelemetry spans through the SDK: `span()` blocks, tool `.action()`s and
LLM calls, with the ids, parents and timings the records carry, so a
collector shows exactly the tree that was logged. Spans are named and
attributed per the OpenTelemetry GenAI conventions (`invoke_agent`,
`execute_tool`, `chat`, with token counts, model and finish reason).
- `OTelLogsTransport` sends every record as an OTLP log record, with the same
trace and span ids so logs and spans join up.
- The convention names live in one file, `logquill/semconv.py`, pinned to a
named release (semantic-conventions 1.44.0), because the GenAI conventions
are still experimental and have already renamed attributes.
`OTEL_SEMCONV_STABILITY_OPT_IN` is honored, `semconv_version="legacy"` picks
the older names, and old and new names are never emitted together. Prompt
and completion text is opt-in only.
- `span(name, capture_state=...)` records what changed during a block as
`meta.state_diff`, and a tool `.action()` that is reopened before it
succeeded gets `meta.retry_count` automatically.
- The record contract gained optional `meta` fields for this: `tool`,
`tool_call_id`, `provider`, `agent_name`, `agent_id`, `response_model`,
`operation`, and opt-in `input_messages` / `output_messages`.
- Type-checking no longer depends on whether OpenTelemetry is installed.
- **Breaking: the record shape and the Python floor changed.** See
[MIGRATING.md](MIGRATING.md).
- logquill now requires **Python 3.10 or newer** and is tested on 3.10–3.14.
Expand Down
9 changes: 7 additions & 2 deletions MIGRATING.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,8 +47,10 @@ LogQuill record, or one written by a newer major version.

## 3. New optional fields

These are defined now so both languages agree on them. Nothing emits them on its
own yet, and no record has them unless you add them.
These are defined so both languages agree on them. `logger.llm_call()` writes
the `llm` block, `span(capture_state=...)` writes `state_diff`, and repeated tool
calls get `retry_count` automatically; nothing else writes them, and no record
has them unless you use those.

| Field | Where | Type |
|---|---|---|
Expand All @@ -58,6 +60,9 @@ own yet, and no record has them unless you add them.
| `meta.retry_count` | `meta` | integer ≥ 0 |
| `meta.state_diff` | `meta` | object |
| `meta.mcp.server`, `meta.mcp.tool` | `meta` | string |
| `meta.tool`, `meta.tool_call_id`, `meta.provider`, `meta.agent_name`, `meta.agent_id`, `meta.response_model` | `meta` | string |
| `meta.operation` | `meta` | `chat`, `text_completion` or `invoke_agent` |
| `meta.input_messages`, `meta.output_messages` | `meta` | array (opt-in prompt/completion content) |

`llm` is its own block, not part of `meta`, because cost and latency
dashboards need stable names for it. A record that isn't an LLM call has no
Expand Down
131 changes: 131 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,7 @@ for what's landed so far.
- **Config from file/env** — `load_config(dict)`, `logger_from_file(path)` (JSON/YAML), `logger_from_env()` build a `Logger` from one config shape — see [Config](#config)
- **Plugin pipeline** — `ContextPlugin`, `RedactPlugin` (by key), `PIIRedactPlugin` (by pattern), `SamplingPlugin` (with tail-based elevation), `TamperEvidentPlugin` (hash-chained logs), `TraceContextPlugin` (cross-service trace correlation), and `AlertingPlugin` (`SlackAlertPlugin`/`PagerDutyAlertPlugin`/`EmailAlertPlugin`, deduplicated) out of the box; a broken plugin can't crash logging; `.use()` also accepts a plain function, no subclassing required (see [Plugins](#plugins))
- **Agentic & harness tracing** — `.child()` loggers, `RunPlugin`, `.thought()/.action()/.observation()/.decision()`, `with agent_log.span(...)`, and framework adapters — `LangChainAdapter` (`pip install logquill[langchain]`), `LangGraphAdapter` (`pip install logquill[langgraph]`, adds checkpoint interrupt/resume events on top), `CrewAIAdapter` (`pip install logquill[crewai]`), `LlamaIndexAdapter` (`pip install logquill[llamaindex]`), and `AutoGenAdapter` (`pip install logquill[autogen]`) — see [Agentic & harness tracing](#agentic--harness-tracing)
- **LLM calls & OpenTelemetry** — `logger.llm_call(model=, tokens_in=, tokens_out=, cost_usd=, ...)` records a first-class `llm` block, `span(capture_state=...)` records what state changed, repeated tool calls get `meta.retry_count` automatically, and `OTLPTransport` (`pip install logquill[otel]`) exports it all as real OpenTelemetry spans named and attributed per the GenAI semantic conventions; `OTelLogsTransport` sends records as OTLP logs — see [LLM calls & OpenTelemetry](#llm-calls--opentelemetry)
- **Non-blocking async dispatch** — `Logger(async_dispatch=True)` moves transport writes onto a background thread with a bounded queue and a configurable backpressure policy (`drop_oldest`/`drop_newest`/`block`); `flush()`/`flush_async()` and a `with_lambda`/`with_cloud_function`/`with_azure_function` decorator make serverless shutdown safe — see [Async dispatch & serverless safety](#async-dispatch--serverless-safety)
- **Zero required runtime dependencies** — stdlib only; `aiohttp` (`logquill[http]`) is opt-in, for keep-alive HTTP delivery
- **Typed throughout** — `mypy --strict` clean on the public API
Expand Down Expand Up @@ -817,6 +818,136 @@ event classes; that's a real divergence, not just a detail, so it needs its
own adapter rather than reusing this one. `autogen-core` is never imported
unless you import `logquill.adapters.autogen` yourself.

## LLM calls & OpenTelemetry

**Recording an LLM call.** `logger.llm_call()` writes one record with the LLM
call's numbers in their own `llm` block — fixed names and types, so a cost or
latency dashboard can rely on them:

```python
from logquill import Logger

logger = Logger("app.agent")

record = logger.llm_call(
"chat",
model="example-model",
tokens_in=1200,
tokens_out=340,
cost_usd=0.0123,
latency_ms=2150.5,
finish_reason="stop",
provider="anthropic", # extra keywords go in `meta`
)

assert record["llm"]["tokens_in"] == 1200
assert record["meta"] == {"kind": "action", "provider": "anthropic"}
```

Leave out what you don't have; a value of the wrong type (a negative count, a
string where a number belongs) is dropped with a warning rather than raising.
Don't put prompt or completion text in `meta` unless you mean every transport
to receive it.

**What changed during a step.** `span(name, capture_state=...)` calls your
function on entering and leaving the block and records what differs as
`meta.state_diff` — only the keys that changed, deep-copied so in-place
mutation is seen, and omitted if nothing changed:

```python
from logquill import CollectingTransport, Logger

sink = CollectingTransport()
logger = Logger("app.agent", transports=[sink])
state = {"items": 1, "user": "ada"}

with logger.span("add_item", capture_state=lambda: state):
state["items"] = 2

assert sink.records[0]["meta"]["state_diff"] == {"before": {"items": 1}, "after": {"items": 2}}
```

`state_diff` lands in `meta`, so `PIIRedactPlugin` sees it; `RedactPlugin` only
matches top-level keys, so don't capture secrets.

**Retries.** An `.action()` that names its tool (`tool="search"`) is tracked:
if the same call — same tool, same enclosing span, same `tool_call_id` if you
give one — is reopened before it succeeded, the record gets `meta.retry_count`
(1, 2, 3, ...). A successful `.observation(tool=...)` ends the chain, so a loop
that legitimately calls one tool many times isn't reported as retries:

```python
from logquill import Logger

logger = Logger("app.agent")

first = logger.action("look it up", tool="search")
logger.observation("timed out", tool="search", error="TimeoutError: slow")
second = logger.action("look it up", tool="search")

assert "retry_count" not in first["meta"]
assert second["meta"]["retry_count"] == 1
```

**Exporting to OpenTelemetry.** `pip install logquill[otel]`, then attach
`OTLPTransport`. It turns the records that describe work with a start and an
end into real spans, with the ids, parents and timings the records carry:

- a `span()` marked `operation="invoke_agent"` (or given an `agent_name`) is
`invoke_agent {agent}`; any other span keeps its own name,
- an `.action()` naming its `tool` is `execute_tool {tool}`,
- a record with an `llm` block is `chat {model}` with the model, token counts
and finish reason as the standard GenAI attributes.

```python
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter

from logquill import Logger, OTLPTransport, RunPlugin

exporter = InMemorySpanExporter() # in real use: leave `span_exporter` out and set `endpoint`
logger = Logger(
"app.agent",
transports=[OTLPTransport(span_exporter=exporter, processor="simple")],
plugins=[RunPlugin()],
)

with logger.span("run", operation="invoke_agent", agent_name="planner"):
logger.llm_call("chat", model="example-model", tokens_in=1200, tokens_out=340, provider="anthropic")
logger.action("look it up", tool="search", duration_ms=40)
logger.close()

spans = {span.name: span for span in exporter.get_finished_spans()}
assert set(spans) == {"invoke_agent planner", "chat example-model", "execute_tool search"}
chat = spans["chat example-model"]
assert chat.attributes["gen_ai.usage.input_tokens"] == 1200
assert chat.parent.span_id == spans["invoke_agent planner"].context.span_id
```

To send to a collector, give it an endpoint instead:
`OTLPTransport(endpoint="http://localhost:4318/v1/traces", service_name="my-agent")`.
Export is batched off the calling thread. Records that aren't spans are
ignored here; `OTelLogsTransport(endpoint=".../v1/logs")` sends every record
as an OTLP log record — level as severity, `meta` as attributes, and the same
trace and span ids, so a collector can join a log line to its span.

Three things worth knowing:

- **The attribute names are pinned to one release of the GenAI conventions**
(semantic-conventions 1.44.0) and live in a single file,
`logquill/semconv.py`. Those conventions are still marked *Development*
upstream and have already renamed attributes (`gen_ai.system` became
`gen_ai.provider.name`), so a future release is an edit to that file.
`semconv_version="legacy"` selects the older provider and token-count names
for a backend that hasn't caught up; `OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental`
selects the latest. Old and new names are never emitted together.
- **Prompt and completion text is never exported by default.** It's only
included — from `meta.input_messages` / `meta.output_messages`, as JSON — if
you pass `capture_content=True` or set
`OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true`.
- Give records a `run_id` (`RunPlugin`) or a `trace_id` (`TraceContextPlugin`)
and one run is one trace. Without either, spans that name a parent still share
a trace with their siblings, but not with their grandparents.

## Async dispatch & serverless safety

By default, every log call dispatches to its transports synchronously — a
Expand Down
4 changes: 4 additions & 0 deletions logquill/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,8 @@
from logquill.transports.nosql.dynamodb_transport import DynamoDBTransport
from logquill.transports.nosql.mongodb_transport import MongoDBTransport
from logquill.transports.nosql.redis_transport import RedisTransport
from logquill.transports.otel.logs_transport import OTelLogsTransport
from logquill.transports.otel.otlp_transport import OTLPTransport
from logquill.transports.queue.base_queue_transport import BaseQueueTransport
from logquill.transports.queue.kafka_transport import KafkaTransport
from logquill.transports.queue.pubsub_transport import PubSubTransport
Expand Down Expand Up @@ -86,6 +88,8 @@
"MongoDBTransport",
"MySQLTransport",
"NewRelicTransport",
"OTLPTransport",
"OTelLogsTransport",
"OptLogger",
"PIIRedactPlugin",
"PagerDutyAlertPlugin",
Expand Down
82 changes: 78 additions & 4 deletions logquill/logger.py
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

import contextlib
import logging
from typing import Any
from typing import Any, Callable

from logquill import shutdown
from logquill.context import current_context
Expand All @@ -11,7 +11,8 @@
from logquill.opt import OptLogger, caller_info, resolve_lazy
from logquill.plugins.context_plugin import ContextPlugin
from logquill.plugins.plugin import FunctionPlugin, MiddlewareFunc, Plugin
from logquill.records import LogRecord, create_record
from logquill.records import LLMBlock, LogRecord, build_llm_block, create_record
from logquill.retry import RetryTracker
from logquill.span import SpanContext, current_span_id
from logquill.toggle import is_enabled
from logquill.transports.transport import Transport
Expand Down Expand Up @@ -72,6 +73,7 @@ def __init__(
if async_dispatch
else None
)
self._retries = RetryTracker()
if flush_at_exit:
shutdown.register(self)

Expand Down Expand Up @@ -166,6 +168,24 @@ def opt(self, *, lazy: bool = False, depth: int | None = None) -> OptLogger:
"""
return OptLogger(self, lazy=lazy, depth=depth)

def _track_retries(self, level: Level, meta: dict[str, Any]) -> None:
"""Stamp `meta.retry_count` on a tool `.action()` that reopens a call
which hasn't succeeded yet, and end the chain on a successful
`.observation()` for it. Only records naming their tool in `meta.tool`
take part."""
tool = meta.get("tool")
if not isinstance(tool, str) or not tool:
return
call_id = meta.get("tool_call_id")
call_id = call_id if isinstance(call_id, str) else None
kind = meta.get("kind")
if kind == "action":
count = self._retries.opened(current_span_id(), tool, call_id)
if count > 0:
meta.setdefault("retry_count", count)
elif kind == "observation" and level < Level.ERROR and "error" not in meta:
self._retries.succeeded(current_span_id(), tool, call_id)

def _local_redactor(self) -> LocalRedactor:
"""Chain every plugin's `redact_local` into one function for
`diagnose` mode. A plugin whose hook raises fails closed: the value
Expand Down Expand Up @@ -216,13 +236,16 @@ def _log(
*,
lazy: bool = False,
depth: int | None = None,
llm: LLMBlock | None = None,
) -> LogRecord | None:
if level < self._level or not is_enabled(self.name):
return None

if lazy:
meta = resolve_lazy(meta)

self._track_retries(level, meta)

if depth is not None:
caller = caller_info(depth)
if caller is not None:
Expand Down Expand Up @@ -250,7 +273,7 @@ def _log(
if stack is not None:
meta["stack"] = stack

record = create_record(level=level, logger=self.name, message=message, meta=meta)
record = create_record(level=level, logger=self.name, message=message, meta=meta, llm=llm)

bound_context = current_context()
if bound_context:
Expand Down Expand Up @@ -347,13 +370,48 @@ def decision(self, message: str, /, **meta: Any) -> LogRecord | None:
decision for a step or run, for harness/agentic tracing."""
return self._log(Level.INFO, message, {"kind": "decision", **meta})

def llm_call(
self,
message: str = "llm_call",
/,
*,
model: str | None = None,
tokens_in: int | None = None,
tokens_out: int | None = None,
cost_usd: float | None = None,
latency_ms: float | None = None,
finish_reason: str | None = None,
**meta: Any,
) -> LogRecord | None:
"""Log one LLM call as an `.action()` carrying the record's first-class
`llm` block (`model`, `tokens_in`, `tokens_out`, `cost_usd`,
`latency_ms`, `finish_reason`). Those fields are what cost and latency
dashboards read, and what `OTLPTransport` exports as the standard token
and model attributes. Any you leave out are simply absent; one of the
wrong type is dropped with a warning rather than raising.

Extra keyword arguments go in `meta` as usual — e.g. `provider="openai"`.
Don't put prompt or completion text in `meta` unless you mean every
transport to receive it.
"""
llm = build_llm_block(
model=model,
tokens_in=tokens_in,
tokens_out=tokens_out,
cost_usd=cost_usd,
latency_ms=latency_ms,
finish_reason=finish_reason,
)
return self._log(Level.INFO, message, {"kind": "action", **meta}, llm=llm)

def span(
self,
name: str,
/,
*,
span_id: str | None = None,
parent_span_id: str | None = None,
capture_state: Callable[[], Any] | None = None,
**meta: Any,
) -> SpanContext:
"""`with agent_log.span("call_llm"):` — on exit, emits one record
Expand All @@ -368,5 +426,21 @@ def span(
`span_id`/`parent_span_id` normally auto-generate/auto-nest; pass
them explicitly to adopt an id handed in from elsewhere (see
`logquill.adapters.langchain.LangChainAdapter` for an example).

`capture_state=lambda: {...}` is called on entering and on leaving the
block, and what changed between the two is recorded as
`meta.state_diff` (`{"before": ..., "after": ...}`, only the keys that
changed when both are dicts; omitted if nothing did). The values are
deep-copied, so an in-place mutation is seen; a state that can't be
copied, or a callable that raises, just means no `state_diff`. It's
stored in `meta`, so a redaction plugin that recurses (`PIIRedactPlugin`)
sees it, but `RedactPlugin` only matches top-level keys.
"""
return SpanContext(self, name, span_id=span_id, parent_span_id=parent_span_id, **meta)
return SpanContext(
self,
name,
span_id=span_id,
parent_span_id=parent_span_id,
capture_state=capture_state,
**meta,
)
Loading
Loading