Skip to content

feat: Add ODBC as a database source (input adapter) #251

Description

@malon64

Summary

Add ODBC as a new input source format, enabling Floe to read data directly from any ODBC-compliant database instead of raw files.

Motivation

Floe currently supports file-based sources (CSV, Parquet, JSON, XLSX, Avro, XML, ORC). Many data platforms source data from relational databases — PostgreSQL, MySQL, SQL Server, Oracle, Snowflake, Redshift, BigQuery — before any file export step. An ODBC source adapter would allow Floe to act as a quality gate directly on database query results, removing the need for an intermediate file stage.

ODBC is the right abstraction here: it is a universal C-level API with drivers available for virtually every database, and the arrow-odbc Rust crate reads ODBC result sets directly into Arrow RecordBatch — which maps immediately to Polars DataFrame with no intermediate conversion.

JDBC is explicitly out of scope: it is a Java API requiring JNI + a live JVM, incompatible with Floe's native Rust model.

Technical Feasibility

The existing InputAdapter trait (io/format.rs:92) is the correct extension point. The only architectural adaptation needed is that a DB source has no file list — it represents a single query result. This is handled by treating the query result as a single virtual InputFile with a synthetic identifier (e.g. odbc://dsn/query), so the rest of the pipeline (validate → split → sink) is completely unchanged.

The Arrow-native path via arrow-odbc:

use arrow_odbc::OdbcReaderBuilder;
use polars::prelude::*;

let reader = OdbcReaderBuilder::new()
    .with_max_bytes_per_batch(512 * 1024 * 1024)
    .build(connection, &query)?;

let mut frames = vec![];
for batch in reader {
    let batch: RecordBatch = batch?;
    frames.push(DataFrame::try_from(batch)?);
}
let df = concat_df(&frames)?;

No custom type mapping required — Arrow handles the ODBC→Polars type bridge.

Proposed Implementation

1. New dependency

# crates/floe-core/Cargo.toml
arrow-odbc = { version = "...", optional = true }
odbc-api = { version = "...", optional = true }

[features]
odbc = ["arrow-odbc", "odbc-api"]

Making it an optional feature avoids forcing the ODBC driver dependency on users who don't need it.

2. Config extension

Add an odbc: Option<OdbcSourceConfig> block to SourceConfig (config/types.rs):

pub struct OdbcSourceConfig {
    pub connection_string: String,  // ODBC DSN or full connection string
    pub query: String,              // SQL SELECT to execute
    pub fetch_size: Option<usize>,  // Rows per Arrow batch (default: 65535)
}

SourceConfig.path becomes unused for ODBC sources (treated as optional at validation time).

User config:

source:
  format: odbc
  odbc:
    connection_string: "Driver={PostgreSQL Unicode};Server=myhost;Database=mydb;Uid=user;Pwd=secret;"
    query: "SELECT id, name, amount, created_at FROM orders WHERE status = 'pending'"

3. New adapter

Create crates/floe-core/src/io/read/odbc.rs implementing InputAdapter:

  • format() → "odbc"
  • default_globs() / suffixes() → return empty (no file discovery)
  • resolve_local_inputs() → return a single synthetic InputFile representing the query
  • read_input_columns() → execute SELECT ... LIMIT 0 or query information_schema to retrieve column names
  • read_inputs() → execute the SQL query, stream Arrow batches, assemble Polars DataFrame

4. Dispatch registration

// io/format.rs
"odbc" => Ok(io::read::odbc::odbc_input_adapter()),

Credentials Handling

The ODBC connection string contains credentials. Options:

  • Reference an existing Floe storage credentials entry (preferred)
  • Support environment variable substitution in connection_string (e.g. Pwd=${DB_PASSWORD})

Plaintext credentials in config files should be explicitly discouraged in documentation.

Known Constraints

  • ODBC drivers are an ops dependency: unixODBC must be installed on the host (Linux/macOS). This is a system-level requirement, not a Rust dependency. Docker-based deployments will need the appropriate driver in the image.
  • No file-level parallelism: unlike file sources, a DB query produces a single result. Parallelism (if needed) would require query partitioning (e.g. WHERE id BETWEEN ? AND ?) — out of scope for the initial implementation.
  • Read-only: this is a source adapter only. Writing back to a database is handled separately by sink adapters (see feat: Add ClickHouse as a new accepted sink format #250 for ClickHouse).

Out of Scope (for now)

  • JDBC (requires JVM — not viable in Rust)
  • Native per-database drivers (sqlx-based PostgreSQL/MySQL adapters) — potential Phase 2 to remove the ODBC driver ops dependency for common databases
  • Query partitioning / parallel reads
  • Incremental / watermark-based reads
  • Write-back / upsert via ODBC

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions