airflow.providers.common.ai.utils.file_analysis

Helpers for building file-analysis prompts for LLM operators.

Attributes

bz2

lzma

SUPPORTED_FILE_FORMATS

log

Classes

FileAnalysisRequest

Prepared prompt content and discovery metadata for the file-analysis operator.

ColumnarSample

A Parquet or Avro file described for a model.

Functions

build_file_analysis_request(*, file_path, ...)

Resolve files, normalize supported formats, and build prompt content for an LLM run.

detect_compression(path)

Return the codec a path's last suffix names, if this Python build can decompress it.

detect_file_format(path)

Detect the logical file format and compression codec from a path suffix.

sample_columnar_file(path, *, file_format, ...)

Describe a Parquet or Avro file for a model: its schema, its first sample_rows rows and its row count.

read_bytes(path, *, compression, max_bytes)

Read path, decompressing it with compression, and refuse more than max_bytes.

Module Contents

airflow.providers.common.ai.utils.file_analysis.bz2 = None[source]
airflow.providers.common.ai.utils.file_analysis.lzma = None[source]
airflow.providers.common.ai.utils.file_analysis.SUPPORTED_FILE_FORMATS: tuple[str, ...] = ('avro', 'csv', 'jpeg', 'jpg', 'json', 'log', 'md', 'parquet', 'pdf', 'png', 'txt')[source]
airflow.providers.common.ai.utils.file_analysis.log[source]
class airflow.providers.common.ai.utils.file_analysis.FileAnalysisRequest[source]

Prepared prompt content and discovery metadata for the file-analysis operator.

user_content: str | collections.abc.Sequence[pydantic_ai.messages.UserContent][source]
resolved_paths: list[str][source]
total_size_bytes: int[source]
omitted_files: int = 0[source]
text_truncated: bool = False[source]
attachment_count: int = 0[source]
text_file_count: int = 0[source]
class airflow.providers.common.ai.utils.file_analysis.ColumnarSample[source]

A Parquet or Avro file described for a model.

text: str[source]

Its schema and first rows.

total_rows: int[source]

How many rows the whole file holds.

airflow.providers.common.ai.utils.file_analysis.build_file_analysis_request(*, file_path, file_conn_id, prompt, multi_modal, max_files, max_file_size_bytes, max_total_size_bytes, max_text_chars, sample_rows)[source]

Resolve files, normalize supported formats, and build prompt content for an LLM run.

airflow.providers.common.ai.utils.file_analysis.detect_compression(path)[source]

Return the codec a path’s last suffix names, if this Python build can decompress it.

airflow.providers.common.ai.utils.file_analysis.detect_file_format(path)[source]

Detect the logical file format and compression codec from a path suffix.

airflow.providers.common.ai.utils.file_analysis.sample_columnar_file(path, *, file_format, sample_rows, max_bytes)[source]

Describe a Parquet or Avro file for a model: its schema, its first sample_rows rows and its row count.

Raises:

LLMFileAnalysisLimitExceededError – if the file is larger than max_bytes.

airflow.providers.common.ai.utils.file_analysis.read_bytes(path, *, compression, max_bytes)[source]

Read path, decompressing it with compression, and refuse more than max_bytes.

Raises:

LLMFileAnalysisLimitExceededError – if the content is larger than max_bytes.

Was this entry helpful?