airflow.providers.common.ai.utils.query_results¶
Bounded payloads for the query and get_schema tools of the SQL toolsets.
A tool result stays in the model’s message history for the rest of the run, so its cost is re-paid on every subsequent request. Three things here keep that bounded:
Columnar query results.
{"columns": [...], "rows": [[...], ...]}names each column once instead of repeating it in a dict per row. On a table with thousands of columns the repeated names, not the values, are the bulk of the payload.A byte budget.
max_rowsandmax_columnscap how many rows or columns come back, which says nothing about size – a single row of a 3000-column table dwarfs a thousand rows of a narrow one. The budget here is what actually bounds context, and when it bites the payload says so, in terms the agent can act on (narrow the projection, or filter the columns).A column cap for get_schema. Column names are what the agent needs to write SQL, so above
max_columnsthe full list is replaced by a summary (count, a type histogram, a sample) that points at thename_containsfilter rather than truncating blindly.
Attributes¶
Functions¶
|
Render query rows as a bounded, columnar JSON tool result. |
|
Render a table's columns as a bounded JSON tool result. |
Module Contents¶
- airflow.providers.common.ai.utils.query_results.QUERY_TOOL_DESCRIPTION = 'Execute a SQL query. Returns JSON of the form {"columns": [name, ...], "rows": [[value, ...],...[source]¶
- airflow.providers.common.ai.utils.query_results.GET_SCHEMA_TOOL_DESCRIPTION = 'Get a table\'s columns. Returns JSON of the form {"columns": [{"name": ..., "type": ...}, ...],...[source]¶
- airflow.providers.common.ai.utils.query_results.build_query_result(columns, rows, *, max_rows, max_result_bytes, more_rows_available, total_rows=None)[source]¶
Render query rows as a bounded, columnar JSON tool result.
- Parameters:
columns (collections.abc.Sequence[str]) – Column names, in the order the values appear in each row.
rows (collections.abc.Sequence[collections.abc.Sequence[Any]]) – Rows already capped to
max_rows; only the byte budget is applied here.max_rows (int) – The row cap that produced rows, reported back to the agent so it knows which limit it hit.
max_result_bytes (int) – Budget for the serialized column names plus rows. The surrounding envelope (the
truncated/hintkeys) adds a small fixed amount on top.more_rows_available (bool) – Whether the query matched more rows than rows holds.
total_rows (int | None) – Total rows the driver reported for the query, when it reports one at all.
Noneis common and is not an error – SQLite and several warehouse drivers do not populate it forSELECT.
- airflow.providers.common.ai.utils.query_results.build_schema_result(columns, *, max_columns, max_result_bytes, name_contains=None)[source]¶
Render a table’s columns as a bounded JSON tool result.
Column names are the information an agent needs to write SQL, so unlike query rows they cannot simply be dropped: above
max_columns(or the byte budget) the full list is replaced by a summary – count, a type histogram, and a sample of columns – that namesname_containsas the way to retrieve specific columns.columnsis assumed to hold distinct names (a table’s introspected columns are unique by construction).- Parameters:
columns (collections.abc.Sequence[dict[str, str]]) –
{"name", "type"}dicts in table order.max_columns (int) – Column count above which a summary replaces the full list.
max_result_bytes (int) – Budget for the serialized result.
name_contains (str | None) – Case-insensitive substring. When given (and non-empty), only matching columns are considered and the value is echoed back so a filtered subset is never mistaken for the whole table.