DataScanDiff
The result of comparing two DataScan profiles.
Usage
DataScanDiff(
current,
baseline,
)Created by calling DataScan.compare(). Provides programmatic access to schema changes and per-column statistical drift, plus a tabular report via get_tabular_report().
Attributes
columns_added: list[str]-
Column names present in the current scan but not the baseline.
columns_removed: list[str]-
Column names present in the baseline but not the current scan.
columns_type_changed: list[str]-
Column names whose data type changed between baseline and current.
column_diffs: list[_ColumnDiff]- Per-column diff details for all columns that appear in either scan.
Attributes
| Name | Description |
|---|---|
| has_changes |
Return True if any schema or statistical changes were detected.
|
| row_count_diff |
Return (baseline_row_count, current_row_count).
|
has_changes
Return True if any schema or statistical changes were detected.
has_changes: bool
row_count_diff
Return (baseline_row_count, current_row_count).
row_count_diff: tuple[int, int]
Methods
| Name | Description |
|---|---|
| get_tabular_report() | Generate a GT table summarizing the differences between the two scans. |
| to_dict() | Export the comparison results as a dictionary. |
get_tabular_report()
Generate a GT table summarizing the differences between the two scans.
Usage
get_tabular_report()Returns
GT- A styled Great Tables report showing schema and statistical drift.
to_dict()
Export the comparison results as a dictionary.
Usage
to_dict()Returns
dict[str, Any]- A dictionary with schema changes, row count diff, and per-column stat diffs.