You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Add ODBC as a new input source format, enabling Floe to read data directly from any ODBC-compliant database instead of raw files.
Motivation
Floe currently supports file-based sources (CSV, Parquet, JSON, XLSX, Avro, XML, ORC). Many data platforms source data from relational databases — PostgreSQL, MySQL, SQL Server, Oracle, Snowflake, Redshift, BigQuery — before any file export step. An ODBC source adapter would allow Floe to act as a quality gate directly on database query results, removing the need for an intermediate file stage.
ODBC is the right abstraction here: it is a universal C-level API with drivers available for virtually every database, and the arrow-odbc Rust crate reads ODBC result sets directly into Arrow RecordBatch — which maps immediately to Polars DataFrame with no intermediate conversion.
JDBC is explicitly out of scope: it is a Java API requiring JNI + a live JVM, incompatible with Floe's native Rust model.
Technical Feasibility
The existing InputAdapter trait (io/format.rs:92) is the correct extension point. The only architectural adaptation needed is that a DB source has no file list — it represents a single query result. This is handled by treating the query result as a single virtual InputFile with a synthetic identifier (e.g. odbc://dsn/query), so the rest of the pipeline (validate → split → sink) is completely unchanged.
The Arrow-native path via arrow-odbc:
use arrow_odbc::OdbcReaderBuilder;use polars::prelude::*;let reader = OdbcReaderBuilder::new().with_max_bytes_per_batch(512*1024*1024).build(connection,&query)?;letmut frames = vec![];for batch in reader {let batch:RecordBatch = batch?;
frames.push(DataFrame::try_from(batch)?);}let df = concat_df(&frames)?;
No custom type mapping required — Arrow handles the ODBC→Polars type bridge.
Making it an optional feature avoids forcing the ODBC driver dependency on users who don't need it.
2. Config extension
Add an odbc: Option<OdbcSourceConfig> block to SourceConfig (config/types.rs):
pubstructOdbcSourceConfig{pubconnection_string:String,// ODBC DSN or full connection stringpubquery:String,// SQL SELECT to executepubfetch_size:Option<usize>,// Rows per Arrow batch (default: 65535)}
SourceConfig.path becomes unused for ODBC sources (treated as optional at validation time).
User config:
source:
format: odbcodbc:
connection_string: "Driver={PostgreSQL Unicode};Server=myhost;Database=mydb;Uid=user;Pwd=secret;"query: "SELECT id, name, amount, created_at FROM orders WHERE status = 'pending'"
The ODBC connection string contains credentials. Options:
Reference an existing Floe storage credentials entry (preferred)
Support environment variable substitution in connection_string (e.g. Pwd=${DB_PASSWORD})
Plaintext credentials in config files should be explicitly discouraged in documentation.
Known Constraints
ODBC drivers are an ops dependency: unixODBC must be installed on the host (Linux/macOS). This is a system-level requirement, not a Rust dependency. Docker-based deployments will need the appropriate driver in the image.
No file-level parallelism: unlike file sources, a DB query produces a single result. Parallelism (if needed) would require query partitioning (e.g. WHERE id BETWEEN ? AND ?) — out of scope for the initial implementation.
Summary
Add ODBC as a new input source format, enabling Floe to read data directly from any ODBC-compliant database instead of raw files.
Motivation
Floe currently supports file-based sources (CSV, Parquet, JSON, XLSX, Avro, XML, ORC). Many data platforms source data from relational databases — PostgreSQL, MySQL, SQL Server, Oracle, Snowflake, Redshift, BigQuery — before any file export step. An ODBC source adapter would allow Floe to act as a quality gate directly on database query results, removing the need for an intermediate file stage.
ODBC is the right abstraction here: it is a universal C-level API with drivers available for virtually every database, and the
arrow-odbcRust crate reads ODBC result sets directly into ArrowRecordBatch— which maps immediately to Polars DataFrame with no intermediate conversion.JDBC is explicitly out of scope: it is a Java API requiring JNI + a live JVM, incompatible with Floe's native Rust model.
Technical Feasibility
The existing
InputAdaptertrait (io/format.rs:92) is the correct extension point. The only architectural adaptation needed is that a DB source has no file list — it represents a single query result. This is handled by treating the query result as a single virtual InputFile with a synthetic identifier (e.g.odbc://dsn/query), so the rest of the pipeline (validate → split → sink) is completely unchanged.The Arrow-native path via
arrow-odbc:No custom type mapping required — Arrow handles the ODBC→Polars type bridge.
Proposed Implementation
1. New dependency
Making it an optional feature avoids forcing the ODBC driver dependency on users who don't need it.
2. Config extension
Add an
odbc: Option<OdbcSourceConfig>block toSourceConfig(config/types.rs):SourceConfig.pathbecomes unused for ODBC sources (treated as optional at validation time).User config:
3. New adapter
Create
crates/floe-core/src/io/read/odbc.rsimplementingInputAdapter:format()→"odbc"default_globs()/suffixes()→ return empty (no file discovery)resolve_local_inputs()→ return a single syntheticInputFilerepresenting the queryread_input_columns()→ executeSELECT ... LIMIT 0or queryinformation_schemato retrieve column namesread_inputs()→ execute the SQL query, stream Arrow batches, assemble Polars DataFrame4. Dispatch registration
Credentials Handling
The ODBC connection string contains credentials. Options:
connection_string(e.g.Pwd=${DB_PASSWORD})Plaintext credentials in config files should be explicitly discouraged in documentation.
Known Constraints
unixODBCmust be installed on the host (Linux/macOS). This is a system-level requirement, not a Rust dependency. Docker-based deployments will need the appropriate driver in the image.WHERE id BETWEEN ? AND ?) — out of scope for the initial implementation.Out of Scope (for now)