Streaming read, convert, and inspect 100+ formats. Install iterabledata[mcp] and run iterable-mcp.
io.github.datenoio/iterabledata MCP Server
This MCP server integrates the iterabledata Python library to support streaming read, conversion, and inspection of data files across many formats. It exposes a consistent, iterator-based interface for row-by-row processing, similar to csv.DictReader, covering more than basic CSV workflows.
🛠️ Key Features
Python library interface for reading and writing files row by row
Unified API across many formats (CSV, JSON, Parquet, XML, etc.)
Iterator-based “consistent” access pattern for processing and conversion
Preserves complex nested data structures (as stated in the excerpt)
🚀 Use Cases
Streaming inspection of datasets without loading entire files
Converting between supported formats while keeping nested structures
⚡ Developer Benefits
Consistent iterator-driven workflow for multiple file types
Unified handling for complex nested data structures
⚠️ Limitations
The provided excerpt describes behavior at the library level; specific MCP tools and capabilities are not detailed here.
Iterable Data is a Python library for reading and writing data files row by row in a consistent, iterator-based interface. It provides a unified API for working with various data formats (CSV, JSON, Parquet, XML, etc.) similar to csv.DictReader but supporting many more formats.
This library simplifies data processing and conversion between formats while preserving complex nested data structures (unlike pandas DataFrames which require flattening).
Features
Unified API: Single interface for reading/writing multiple data formats
Automatic Format Detection: Detects file type and compression from filename or content (magic numbers and heuristics)
Format Capability Reporting: Programmatically query format capabilities (read/write/bulk/totals/streaming/tables)
Support for Compression: Works seamlessly with compressed files
Preserves Nested Data: Handles complex nested structures as Python dictionaries
DuckDB Integration: Optional DuckDB engine for high-performance queries with pushdown optimizations
Pipeline Processing: Built-in pipeline support for data transformation
Encoding Detection: Automatic encoding and delimiter detection for text files
Bulk Operations: Efficient batch reading and writing
Native Batch Conversion: Opt-in columnar-to-columnar transfers with projection, row-range, and batch-size selection
Bounded Columnar I/O: Shared row/bulk cursors and configurable Parquet row groups keep large reads and writes bounded
Codec Performance Profiles: Choose fast, balanced, or max compression settings with effective-setting diagnostics
Table Listing: Discover available tables, sheets, and datasets in multi-table formats
Context Manager Support: Use with statements for automatic resource cleanup
DataFrame Bridges: Convert iterable data to Pandas, Polars, and Dask DataFrames with one-liner methods
Cloud Storage Support: Direct access to S3, GCS, and Azure Blob Storage via URI schemes
Database Engine Support: Read-only access to SQL and NoSQL databases (PostgreSQL, ClickHouse, MySQL, MongoDB, Elasticsearch, etc.) as iterable data sources
Atomic Writes: Production-safe file writing with temporary files and atomic renames
Bulk File Conversion: Convert multiple files at once using glob patterns or directories
Progress Tracking and Metrics: Built-in progress bars, callbacks, and structured metrics objects
Error Handling Controls: Configurable error policies and structured error logging; malformed input raises typed errors by default instead of reading as empty datasets
Security Hardening: XXE-safe XML parsing, AST-whitelisted filter expressions, and explicit pickle trust acknowledgement
Performance Regression Gate: CI-enforced baselines for representative read/convert workloads
Container Formats: Stream records from TAR archives without extracting members to disk
Type Hints and Type Safety: Complete type annotations with typed helper functions for dataclasses and Pydantic models
Lakehouse Tables: Read and write Delta Lake, Iceberg, and DuckLake; read Hudi; experimental Apache Paimon tables plus Row/Mosaic file formats
Supported File Types
Core Formats
JSON - Standard JSON files
JSONL/NDJSON - JSON Lines format (one JSON object per line)
JSON-LD - JSON for Linking Data (RDF format)
CSV/TSV - Comma and tab-separated values
Annotated CSV - CSV with type annotations and metadata
CSVW - CSV on the Web (with metadata)
PSV/SSV - Pipe and semicolon-separated values
LTSV - Labeled Tab-Separated Values
FWF - Fixed Width Format
XML - XML files with configurable tag parsing
ZIP XML - XML files within ZIP archives
HTML - HTML files with table extraction
Binary Formats
BSON - Binary JSON format
MessagePack - Efficient binary serialization
CBOR - Concise Binary Object Representation
UBJSON - Universal Binary JSON
SMILE - Binary JSON variant
Bencode - BitTorrent encoding format
Avro - Apache Avro binary format (read & write)
Pickle - Python pickle format (untrusted input is unsafe; pass trust=True to acknowledge)
Columnar & Analytics Formats
Parquet - Apache Parquet columnar format
ORC - Optimized Row Columnar format
Arrow/Feather - Apache Arrow columnar format
GeoParquet - GeoParquet metadata-aware Parquet profile with geometry/CRS preservation
Lance - Modern columnar format optimized for ML and vector search
Vortex - Modern columnar format with fast random access
Paimon Row - Apache Paimon row format for O(1) row-number access
Paimon Mosaic - Apache Paimon columnar-bucket format for wide tables
AI Features ([ai]): Enables AI-powered documentation generation using OpenAI, OpenRouter, Ollama, LMStudio, or Perplexity.
Database Engines ([db]): Enables read-only database access as iterable data sources. Supports PostgreSQL, ClickHouse, MySQL/MariaDB, Microsoft SQL Server, SQLite, MongoDB, and Elasticsearch/OpenSearch. Includes convenience groups:
[db-sql]: SQL databases only (PostgreSQL, ClickHouse, MySQL, MSSQL)
[db-nosql]: NoSQL databases only (MongoDB, Elasticsearch)
Genomic formats ([bio]): Enables genomic VCF/BCF, CRAM, BED, GFF3, and GTF support. Alignment formats use pysam and may require a reference file.
Geospatial / scientific extras: [geospatial] covers FileGDB and MapInfo MIF (plus existing GeoPackage/Shapefile stack). [lidar], [mat], and [geophysical] enable LAS, MATLAB MAT, and SEG-Y/GRIB2/miniSEED respectively. Many structure formats (XYZ, CIF, PDB, ASCII Grid, CZML, EDI, WebDataset, Lotus WK1) need no extra.
Lakehouse ([lakehouse]): Delta Lake, Apache Iceberg, Lance, Apache Hudi, and DuckLake. Delta, Iceberg, and DuckLake support bounded writes; Hudi is read-only for now.
Paimon ([paimon]): Apache Paimon warehouse tables plus Row and Mosaic file formats. Install [paimon-table], [paimon-row], or [paimon-mosaic] individually if you only need one surface. DuckLake alone is also available as [ducklake].
See the API documentation for details on these features.
Generate dataset documentation with a local LLM (no API key) via LM Studio or Ollama, or use OpenAI:
python
from iterable.ai import doc
# Local (LM Studio on http://localhost:1234/v1)
documentation = doc.generate(
"data.csv",
provider="lmstudio",
base_url="http://localhost:1234/v1",
format="markdown",
)
# Or analyze structure + docs in one callfrom iterable.ops import inspect
analysis = inspect.analyze("data.csv", autodoc=True, autodoc_provider="openai")
print(analysis["documentation"])
Need structured, machine-readable output? Use generate_blocks() to get independent documentation blocks (general, schema, quality, examples, statistics, agent_skill; plus opt-in codebook) plus the assembled markdown:
The agent_skill block emits a portable skill document (YAML frontmatter + Markdown) that AI agents can load for dataset-specific load/query/safety guidance.
For multi-table workbooks/databases, iterable.ai.table_profile.profile_selected_table()
profiles one named sheet/table under row/time budgets (nested flattening enabled).
Format catalog (agents)
python
from iterable.catalog import describe_format
info = describe_format("xml")
print(info["example_args"]) # {'tagname': 'item'}
from iterable import open_iterable
with open_iterable("data.csv.gz") as source:
for row in source:
print(row)
Writing Data
python
from iterable import open_iterable
with open_iterable("output.jsonl.zst", mode="w") as dest:
for item in my_data:
dest.write(item)
Usage Examples
Reading Compressed CSV Files
python
from iterable import open_iterable
with open_iterable("data.csv.xz") as source:
n = 0for row in source:
n += 1if n % 1000 == 0:
print(f"Processed {n} rows")
Reading Different Formats
python
from iterable import open_iterable
with open_iterable("data.jsonl") as source:
for row in source:
print(row)
with open_iterable("data.parquet") as source:
for row in source:
print(row)
with open_iterable("data.xml", iterableargs={"tagname": "item"}) as source:
for row in source:
print(row)
with open_iterable("data.xlsx") as source:
for row in source:
print(row)
# Read GeoJSON Text Sequence (streaming, one feature per line)with open_iterable('features.geojsonl') as source:
for feature in source:
print(feature['properties'], feature['geometry'])
# Stream records from a TAR archive (members detected by filename)with open_iterable('dataset.tar.gz', iterableargs={'members': '*.csv'}) as source:
for row in source:
print(row['_member'], row)
# Read genomic VCF (requires pip install iterabledata[bio])with open_iterable('variants.vcf') as source:
for variant in source:
print(variant['chrom'], variant['pos'], variant['ref'], variant['alt'])
# Read genomic intervals (BED, GFF3, or GTF)with open_iterable('genes.gff3') as source:
for feature in source:
print(feature['seqid'], feature['type'], feature['start'], feature['end'])
# Read a Zarr array (requires pip install iterabledata[zarr])with open_iterable('signals.zarr', iterableargs={'array': 'values'}) as source:
for row in source:
print(row['value'])
# Read GeoParquet or FlatGeobuf (requires parquet/geospatial extras)with open_iterable('roads.geoparquet') as source:
for feature in source:
print(feature.get('geometry'), feature.get('properties'))
# Read an OTLP JSON export (requires pip install iterabledata[otlp])with open_iterable('telemetry.otlp.json') as source:
for item in source:
print(item['signal'], item['record'])
Reading from Databases
python
from iterable import open_iterable
# Read from PostgreSQL databasewith open_iterable(
'postgresql://user:password@localhost:5432/mydb',
engine='postgres',
iterableargs={'query': 'users'}
) as source:
for row in source:
print(row)
# Read specific columns with filteringwith open_iterable(
'postgresql://localhost/mydb',
engine='postgres',
iterableargs={
'query': 'users',
'columns': ['id', 'name', 'email'],
'filter': 'active = TRUE'
}
) as source:
for row in source:
print(row)
# Read from ClickHouse databasewith open_iterable(
'clickhouse://user:password@localhost:9000/analytics',
engine='clickhouse',
iterableargs={'query': 'events', 'settings': {'max_threads': 4}}
) as source:
for row in source:
print(row)
# Convert database to filefrom iterable.convert import convert
convert(
fromfile='postgresql://localhost/mydb',
tofile='users.parquet',
iterableargs={'engine': 'postgres', 'query': 'users'}
)
# Convert ClickHouse to Parquet
convert(
fromfile='clickhouse://localhost:9000/analytics',
tofile='events.parquet',
iterableargs={'engine': 'clickhouse', 'query': 'events'}
)
Format Detection and Encoding
python
from iterable import open_iterable
from iterable.helpers.detect import detect_file_type, detect_file_type_from_content
from iterable.helpers.utils import detect_encoding, detect_delimiter
# Detect file type and compression (uses filename extension)
result = detect_file_type('data.csv.gz')
print(f"Type: {result['datatype']}, Codec: {result['codec']}")
# Content-based detection (for files without extensions or streams)withopen('data.unknown', 'rb') as f:
detection_result = detect_file_type_from_content(f)
if detection_result:
format_id, confidence, method = detection_result
print(f"Detected format: {format_id} (confidence: {confidence:.2f}, method: {method})")
# open_iterable() automatically uses content-based detection as fallback# Works with files without extensions, streams, or incorrect extensionswith open_iterable('data.unknown') as source: # Detects from contentfor row in source:
print(row)
# Detect encoding for CSV files
encoding_info = detect_encoding('data.csv')
print(f"Encoding: {encoding_info['encoding']}, Confidence: {encoding_info['confidence']}")
# Detect delimiter for CSV files
delimiter = detect_delimiter('data.csv', encoding=encoding_info['encoding'])
# Open with detected settings
source = open_iterable('data.csv', iterableargs={
'encoding': encoding_info['encoding'],
'delimiter': delimiter
})
Error Handling
IterableData provides a comprehensive exception hierarchy and configurable error handling:
python
from iterable import open_iterable
from iterable.exceptions import (
FormatDetectionError,
FormatNotSupportedError,
FormatParseError,
ReadError,
CodecError,
IterableDataError,
)
# Basic exception handlingtry:
with open_iterable('data.unknown') as source:
for row in source:
process(row)
except FormatDetectionError as e:
print(f"Could not detect format: {e.reason}")
# Try with explicit format or check file contentexcept FormatNotSupportedError as e:
print(f"Format '{e.format_id}' not supported: {e.reason}")
# Install missing dependencies or use different formatexcept FormatParseError as e:
print(f"Failed to parse {e.format_id} format")
if e.position:
print(f"Error at position: {e.position}")
except ReadError as e:
print(f"Read failed: {e}")
except IterableDataError as e:
print(f"Library error: {e}")
except CodecError as e:
print(f"Compression error with {e.codec_name}: {e.message}")
# Check file integrity or try different codecexcept Exception as e:
print(f"Unexpected error: {e}")
Configurable Error Policies: Control how malformed records are handled:
python
# Skip malformed records and continue processingwith open_iterable(
'data.csv',
iterableargs={'on_error': 'skip', 'error_log': 'errors.log'}
) as src:
for row in src:
process(row) # Only processes valid rows# Warn on errors but continue processingwith open_iterable(
'data.jsonl',
iterableargs={'on_error': 'warn', 'error_log': 'errors.log'}
) as src:
for row in src:
process(row) # Warnings logged, processing continues# Default: raise exceptions immediately (existing behavior)with open_iterable('data.csv', iterableargs={'on_error': 'raise'}) as src:
for row in src:
process(row)
No silent empty reads: Under the default policy (on_error='raise'), a malformed non-empty file raises FormatParseError rather than yielding zero records. Use on_error='skip' or 'warn' to tolerate bad records explicitly.
Pickle safety: Unpickling executes arbitrary code. Reading pickle files emits a warning unless you pass trust=True:
python
with open_iterable('data.pickle', iterableargs={'trust': True}) as source:
for row in source:
process(row)
Error Logging: Structured JSON logs with context (filename, row number, byte offset, error message, original line).
from iterable.helpers.capabilities import (
get_format_capabilities,
get_capability,
list_all_capabilities
)
# Get all capabilities for a format
caps = get_format_capabilities("csv")
print(f"CSV readable: {caps['readable']}")
print(f"CSV writable: {caps['writable']}")
print(f"CSV supports totals: {caps['totals']}")
print(f"CSV supports tables: {caps['tables']}")
# Query a specific capability
is_writable = get_capability("json", "writable")
has_totals = get_capability("parquet", "totals")
supports_tables = get_capability("xlsx", "tables")
# List capabilities for all formats
all_caps = list_all_capabilities()
for format_id, capabilities in all_caps.items():
if capabilities.get("tables"):
print(f"{format_id} supports multiple tables")
Format Conversion
python
from iterable import open_iterable
from iterable.convert import convert
# Simple format conversion
convert('input.jsonl.gz', 'output.parquet')
# Convert with options
convert(
'input.csv.xz',
'output.jsonl.zst',
iterableargs={'delimiter': ';', 'encoding': 'utf-8'},
batch_size=10000
)
# Convert and flatten nested structures
convert(
'input.jsonl',
'output.csv',
is_flatten=True,
batch_size=50000
)
Atomic Writes for Production Safety
Use atomic writes to ensure output files are never left in a partially written state:
python
from iterable.convert import convert
from iterable.pipeline import pipeline
# Convert with atomic writes (production-safe)
result = convert('input.csv', 'output.parquet', atomic=True)
# Output file only appears when conversion completes successfully# Atomic writes in pipelines
pipeline(
source=source,
destination=destination,
process_func=transform_func,
atomic=True# Ensures destination file is only created on success
)
Benefits: Prevents data corruption from crashes, interruptions, or mid-process failures. Original files are preserved on failure.
Bulk File Conversion
Convert multiple files at once using glob patterns, directories, or file lists:
python
from iterable.convert import bulk_convert
# Convert all CSV files matching glob pattern
result = bulk_convert('data/raw/*.csv.gz', 'data/processed/', to_ext='parquet')
# Convert with custom filename pattern
result = bulk_convert('data/*.csv', 'output/', pattern='{name}.parquet')
# Convert entire directory
result = bulk_convert('data/raw/', 'data/processed/', to_ext='parquet')
# Access resultsprint(f"Converted {result.successful_files}/{result.total_files} files")
print(f"Total rows: {result.total_rows_out}")
print(f"Throughput: {result.throughput:.0f} rows/second")
# Check individual file resultsfor file_result in result.file_results:
if file_result.success:
print(f"✓ {file_result.source_file}: {file_result.result.rows_out} rows")
else:
print(f"✗ {file_result.source_file}: {file_result.error}")
Features: Error resilience (continues if one file fails), aggregated metrics, flexible output naming with placeholders ({name}, {stem}, {ext}).
Progress Tracking and Metrics
Track conversion and pipeline progress with callbacks, progress bars, and structured metrics:
python
from iterable.convert import convert
from iterable.pipeline import pipeline
# Progress callback for conversionsdefprogress_cb(stats):
print(f"Progress: {stats['rows_read']} rows read, "f"{stats['rows_written']} rows written, "f"{stats.get('elapsed', 0):.2f}s elapsed")
# Convert with progress tracking
result = convert(
'input.csv',
'output.parquet',
progress=progress_cb,
show_progress=True# Also shows tqdm progress bar
)
# Access conversion metricsprint(f"Converted {result.rows_out} rows in {result.elapsed_seconds:.2f}s")
print(f"Read {result.bytes_read} bytes, wrote {result.bytes_written} bytes")
# Pipeline with progress and metrics
result = pipeline(
source=source,
destination=destination,
process_func=transform_func,
progress=progress_cb # Progress callback
)
# Access pipeline metrics (supports both attribute and dict access)print(f"Processed {result.rows_processed} rows")
print(f"Throughput: {result.throughput:.0f} rows/second")
print(f"Exceptions: {result.exceptions}")
# Backward compatible: result['rec_count'] also works
from iterable import open_iterable
from iterable.pipeline import pipeline
source = open_iterable('input.parquet')
destination = open_iterable('output.jsonl.xz', mode='w')
deftransform_record(record, state):
"""Transform each record"""# Add processing logic
out = {}
for key in ['name', 'email', 'age']:
if key in record:
out[key] = record[key]
return out
defprogress_callback(stats, state):
"""Called every trigger_on records"""print(f"Processed {stats['rec_count']} records, "f"Duration: {stats.get('duration', 0):.2f}s")
deffinal_callback(stats, state):
"""Called when processing completes"""print(f"Total records: {stats['rec_count']}")
print(f"Total time: {stats['duration']:.2f}s")
result = pipeline(
source=source,
destination=destination,
process_func=transform_record,
trigger_func=progress_callback,
trigger_on=1000,
final_func=final_callback,
start_state={},
atomic=True# Use atomic writes for production safety
)
# Access pipeline metricsprint(f"Throughput: {result.throughput:.0f} rows/second")
source.close()
destination.close()
Manual Format and Codec Usage
python
from iterable.datatypes.jsonl import JSONLinesIterable
from iterable.datatypes.bsonf import BSONIterable
from iterable.codecs.gzipcodec import GZIPCodec
from iterable.codecs.lzmacodec import LZMACodec
# Read gzipped JSONL
read_codec = GZIPCodec('input.jsonl.gz', mode='r', open_it=True)
reader = JSONLinesIterable(codec=read_codec)
# Write LZMA compressed BSON
write_codec = LZMACodec('output.bson.xz', mode='wb', open_it=False)
writer = BSONIterable(codec=write_codec, mode='w')
for row in reader:
writer.write(row)
reader.close()
writer.close()
Cloud Storage Support
Read and write data directly from cloud object storage (S3, GCS, Azure):
python
from iterable import open_iterable
# Read from S3with open_iterable('s3://my-bucket/data/events.csv') as source:
for row in source:
print(row)
# Read compressed file from GCSwith open_iterable('gs://my-bucket/data/events.jsonl.gz') as source:
for row in source:
process(row)
# Write to Azure Blob Storagewith open_iterable(
'az://my-container/output/results.jsonl',
mode='w',
iterableargs={'storage_options': {'connection_string': '...'}}
) as dest:
dest.write({'name': 'Alice', 'age': 30})
dest.write({'name': 'Bob', 'age': 25})
Supported Providers:
Amazon S3: s3:// and s3a:// schemes
Google Cloud Storage: gs:// and gcs:// schemes
Azure Blob Storage: az://, abfs://, and abfss:// schemes
Installation: pip install iterabledata[cloud]
Note: DuckDB engine does not support cloud storage URIs; use engine='internal' (default).
Using DuckDB Engine with Pushdown Optimizations
The DuckDB engine provides high-performance querying with advanced optimizations:
python
from iterable import open_iterable
# Basic DuckDB usage
source = open_iterable('data.csv.gz', engine='duckdb')
total = source.totals() # Fast countingfor row in source:
print(row)
source.close()
# Column projection pushdown (only read specified columns)with open_iterable(
'data.csv',
engine='duckdb',
iterableargs={'columns': ['name', 'age']} # Reduces I/O and memory
) as src:
for row in src:
process(row)
# Filter pushdown (filter at database level)with open_iterable(
'data.csv',
engine='duckdb',
iterableargs={'filter': "age > 18 AND status = 'active'"}
) as src:
for row in src:
process(row)
# Combined column projection and filteringwith open_iterable(
'data.parquet',
engine='duckdb',
iterableargs={
'columns': ['name', 'age', 'email'],
'filter': 'age > 18'
}
) as src:
for row in src:
process(row)
# Direct SQL query supportwith open_iterable(
'data.parquet',
engine='duckdb',
iterableargs={
'query': 'SELECT name, age FROM read_parquet(\'data.parquet\') WHERE age > 18 ORDER BY age DESC LIMIT 100'
}
) as src:
for row in src:
process(row)
from iterable import open_iterable
source = open_iterable('input.jsonl')
destination = open_iterable('output.parquet', mode='w')
# Read and write in batches for better performance
batch = []
for row in source:
batch.append(row)
iflen(batch) >= 10000:
destination.write_bulk(batch)
batch = []
# Write remaining recordsif batch:
destination.write_bulk(batch)
source.close()
destination.close()
Performance-oriented Conversion
Use native batches when both endpoints are Parquet or Arrow/Feather. The
selection is pushed into the columnar reader when supported; otherwise the
conversion safely falls back to the regular row/bulk loop:
from iterable import open_iterable
from iterable.ai.fileinfo import open_table
# Read Excel file (specify sheet or page)
xls_file = open_iterable('data.xlsx', iterableargs={'page': 0})
for row in xls_file:
print(row)
xls_file.close()
# Open a named sheet (uses page index under the hood)
sheet = open_table('data.xlsx', 'Sheet2')
for row in sheet:
print(row)
sheet.close()
XML Processing
python
from iterable import open_iterable
# Parse XML with specific tag name
xml_file = open_iterable(
'data.xml',
iterableargs={
'tagname': 'book',
'prefix_strip': True# Strip XML namespace prefixes
}
)
for item in xml_file:
print(item)
xml_file.close()
DataFrame Bridges
Convert iterable data to Pandas, Polars, or Dask DataFrames:
python
from iterable import open_iterable
# Convert to Pandas DataFramewith open_iterable('data.csv.gz') as source:
df = source.to_pandas()
print(df.head())
# Chunked processing for large fileswith open_iterable('large_data.csv') as source:
for df_chunk in source.to_pandas(chunksize=100_000):
# Process each chunk
result = df_chunk.groupby('category').sum()
process_chunk(result)
# Convert to Polars DataFramewith open_iterable('data.csv.gz') as source:
df = source.to_polars()
print(df.head())
# Convert to Dask DataFrame (single file)with open_iterable('data.csv.gz') as source:
ddf = source.to_dask()
result = ddf.groupby('category').sum().compute()
# Multi-file Dask DataFrame (automatic format detection)from iterable.helpers.bridges import to_dask
ddf = to_dask(['file1.csv', 'file2.jsonl', 'file3.parquet'])
result = ddf.groupby('category').sum().compute()
pip install iterabledata[dataframes] # All DataFrame libraries# Or individually:
pip install pandas
pip install polars
pip install "dask[dataframe]"
Type Hints and Type Safety
IterableData includes complete type annotations and typed helper functions for modern Python development:
python
from iterable import open_iterable
from iterable.helpers.typed import as_dataclasses, as_pydantic
from dataclasses import dataclass
from pydantic import BaseModel
# Type aliases for better code readabilityfrom iterable.types import Row, IterableArgs, CodecArgs
# Convert to dataclasses for type safety@dataclassclassPerson:
name: str
age: int
email: str | None = Nonewith open_iterable('people.csv') as source:
for person in as_dataclasses(source, Person):
# Full IDE autocomplete and type checkingprint(person.name, person.age)
# Convert to Pydantic models with validationclassPersonModel(BaseModel):
name: str
age: int
email: str | None = Nonewith open_iterable('people.jsonl') as source:
for person in as_pydantic(source, PersonModel, validate=True):
# Automatic schema validationprint(person.name, person.age)
# Access as Pydantic model with all features
Benefits:
Complete type annotations across the public API
py.typed marker file enables mypy, pyright, and other type checkers
Typed helpers provide IDE autocomplete and type safety
Pydantic validation catches schema issues early
Installation: pip install iterabledata[pydantic] for Pydantic support
Advanced: Converting Compressed XML to Parquet
python
from iterable.datatypes.xml import XMLIterable
from iterable.datatypes.parquet import ParquetIterable
from iterable.codecs.bz2codec import BZIP2Codec
# Read compressed XML
read_codec = BZIP2Codec('data.xml.bz2', mode='r')
reader = XMLIterable(codec=read_codec, tagname='page')
# Write to Parquet with schema adaptation
writer = ParquetIterable(
'output.parquet',
mode='w',
use_pandas=False,
adapt_schema=True,
batch_size=10000
)
batch = []
for row in reader:
batch.append(row)
iflen(batch) >= 10000:
writer.write_bulk(batch)
batch = []
if batch:
writer.write_bulk(batch)
reader.close()
writer.close()
Agent skill block (agent_skill): Default generate_blocks() now includes a portable agent-skill document (YAML frontmatter + Markdown) with dataset facts, workflow, and safety guidance
Safer usage examples: Examples / legacy autodoc prompts require python/r/sql, SQL against table dataset, and read-only constraints
Schema examples: Nested provider example values are coerced to strings for structured-output validation