airflow.providers.common.ai.utils.query_results

Bounded payloads for the query and get_schema tools of the SQL toolsets.

A tool result stays in the model’s message history for the rest of the run, so its cost is re-paid on every subsequent request. Three things here keep that bounded:

  • Columnar query results. {"columns": [...], "rows": [[...], ...]} names each column once instead of repeating it in a dict per row. On a table with thousands of columns the repeated names, not the values, are the bulk of the payload.

  • A byte budget. max_rows and max_columns cap how many rows or columns come back, which says nothing about size – a single row of a 3000-column table dwarfs a thousand rows of a narrow one. The budget here is what actually bounds context, and when it bites the payload says so, in terms the agent can act on (narrow the projection, or filter the columns).

  • A column cap for get_schema. Column names are what the agent needs to write SQL, so above max_columns the full list is replaced by a summary (count, a type histogram, a sample) that points at the name_contains filter rather than truncating blindly.

Attributes

DEFAULT_MAX_RESULT_BYTES

DEFAULT_MAX_COLUMNS

QUERY_TOOL_DESCRIPTION

GET_SCHEMA_TOOL_DESCRIPTION

Functions

build_query_result(columns, rows, *, max_rows, ...[, ...])

Render query rows as a bounded, columnar JSON tool result.

build_schema_result(columns, *, max_columns, ...[, ...])

Render a table's columns as a bounded JSON tool result.

Module Contents

airflow.providers.common.ai.utils.query_results.DEFAULT_MAX_RESULT_BYTES = 65536[source]
airflow.providers.common.ai.utils.query_results.DEFAULT_MAX_COLUMNS = 100[source]
airflow.providers.common.ai.utils.query_results.QUERY_TOOL_DESCRIPTION = 'Execute a SQL query. Returns JSON of the form {"columns": [name, ...], "rows": [[value, ...],...[source]
airflow.providers.common.ai.utils.query_results.GET_SCHEMA_TOOL_DESCRIPTION = 'Get a table\'s columns. Returns JSON of the form {"columns": [{"name": ..., "type": ...}, ...],...[source]
airflow.providers.common.ai.utils.query_results.build_query_result(columns, rows, *, max_rows, max_result_bytes, more_rows_available, total_rows=None)[source]

Render query rows as a bounded, columnar JSON tool result.

Parameters:
  • columns (collections.abc.Sequence[str]) – Column names, in the order the values appear in each row.

  • rows (collections.abc.Sequence[collections.abc.Sequence[Any]]) – Rows already capped to max_rows; only the byte budget is applied here.

  • max_rows (int) – The row cap that produced rows, reported back to the agent so it knows which limit it hit.

  • max_result_bytes (int) – Budget for the serialized column names plus rows. The surrounding envelope (the truncated/hint keys) adds a small fixed amount on top.

  • more_rows_available (bool) – Whether the query matched more rows than rows holds.

  • total_rows (int | None) – Total rows the driver reported for the query, when it reports one at all. None is common and is not an error – SQLite and several warehouse drivers do not populate it for SELECT.

airflow.providers.common.ai.utils.query_results.build_schema_result(columns, *, max_columns, max_result_bytes, name_contains=None)[source]

Render a table’s columns as a bounded JSON tool result.

Column names are the information an agent needs to write SQL, so unlike query rows they cannot simply be dropped: above max_columns (or the byte budget) the full list is replaced by a summary – count, a type histogram, and a sample of columns – that names name_contains as the way to retrieve specific columns. columns is assumed to hold distinct names (a table’s introspected columns are unique by construction).

Parameters:
  • columns (collections.abc.Sequence[dict[str, str]]) – {"name", "type"} dicts in table order.

  • max_columns (int) – Column count above which a summary replaces the full list.

  • max_result_bytes (int) – Budget for the serialized result.

  • name_contains (str | None) – Case-insensitive substring. When given (and non-empty), only matching columns are considered and the value is echoed back so a filtered subset is never mistaken for the whole table.

Was this entry helpful?