Lance
Docs /Integrations /Apache Spark /Config

Configuration

Spark DSV2 catalog integrates with Lance through Lance Namespace.

Spark SQL Extensions

Lance provides SQL extensions that add additional functionality beyond standard Spark SQL. To enable these extensions, configure your Spark application with:

spark = SparkSession.builder \
    .appName("lance-example") \
    .config("spark.sql.extensions", "org.lance.spark.extensions.LanceSparkSessionExtensions") \
    .getOrCreate()
val spark = SparkSession.builder()
    .appName("lance-example")
    .config("spark.sql.extensions", "org.lance.spark.extensions.LanceSparkSessionExtensions")
    .getOrCreate()
SparkSession spark = SparkSession.builder()
    .appName("lance-example")
    .config("spark.sql.extensions", "org.lance.spark.extensions.LanceSparkSessionExtensions")
    .getOrCreate();
spark-shell \
  --packages org.lance:lance-spark-bundle-3.5_2.12:0.4.0 \
  --conf spark.sql.extensions=org.lance.spark.extensions.LanceSparkSessionExtensions
spark-submit \
  --packages org.lance:lance-spark-bundle-3.5_2.12:0.4.0 \
  --conf spark.sql.extensions=org.lance.spark.extensions.LanceSparkSessionExtensions \
  your-application.jar

Features Requiring Extensions

The following features require the Lance Spark SQL extension to be enabled:

  • VECTOR_SEARCH - Run vector similarity search through Lance namespace execution
  • SEARCH - Run full-text search through Lance namespace execution
  • HYBRID_SEARCH - Combine vector and full-text search with reciprocal rank fusion
  • Full-Text Search - lance_match, lance_match_phrase, and lance_multi_match SQL functions for querying FTS indexes
  • ADD COLUMNS with backfill - Add new columns and backfill existing rows with data
  • UPDATE COLUMNS with backfill - Update existing columns using data from a source
  • OPTIMIZE - Compact table fragments for improved query performance
  • OPTIMIZE INDEX - Incrementally maintain a named index
  • VACUUM - Remove old versions and reclaim storage space

Basic Setup

Configure Spark with the LanceNamespaceSparkCatalog by setting the appropriate Spark catalog implementation and namespace-specific options:

Parameter Type Required Description
spark.sql.catalog.{name} String ✓ Set to org.lance.spark.LanceNamespaceSparkCatalog
spark.sql.catalog.{name}.impl String ✓ Namespace implementation, short name like dir, rest, hive3, glue or full Java implementation class
spark.sql.catalog.{name}.storage.* - ✗ Lance IO storage options. See Lance Object Store Guide for all available options.
spark.sql.catalog.{name}.index_cache_backend String ✗ Registered native index cache backend URI, for example moka://?capacity=6442450944.
spark.sql.catalog.{name}.metadata_cache_backend String ✗ Registered native metadata cache backend URI, for example moka://?capacity=1073741824.
spark.sql.catalog.{name}.single_level_ns Boolean ✗ Enable single-level mode with virtual "default" namespace. Default: false. See Note on Namespace Levels.
spark.sql.catalog.{name}.parent String ✗ Parent prefix for multi-level namespaces. See Note on Namespace Levels.
spark.sql.catalog.{name}.parent_delimiter String ✗ Delimiter for parent prefix (default: .). See Note on Namespace Levels.

Cache Backends

Each Spark catalog owns an isolated Lance Session. Select registered native cache backends for the index and metadata cache independently by setting backend URIs:

spark = SparkSession.builder \
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog") \
    .config("spark.sql.catalog.lance.impl", "dir") \
    .config("spark.sql.catalog.lance.root", "/path/to/lance/database") \
    .config(
        "spark.sql.catalog.lance.index_cache_backend",
        "moka://?capacity=6442450944",
    ) \
    .config(
        "spark.sql.catalog.lance.metadata_cache_backend",
        "moka://?capacity=1073741824",
    ) \
    .getOrCreate()

moka is registered by Lance Core. Other backend kinds must already be registered by the native Lance runtime packaged with the application. These options select a registered backend; they do not register a Java cache implementation.

The connector serializes the backend URIs with its read and write options, so driver and executor JVMs create equivalent process-local sessions. A catalog's session is fixed when that catalog is first used. To switch backends in one Spark application, configure a second catalog:

spark = SparkSession.builder \
    .config("spark.sql.catalog.memory_lance", "org.lance.spark.LanceNamespaceSparkCatalog") \
    .config("spark.sql.catalog.memory_lance.impl", "dir") \
    .config("spark.sql.catalog.memory_lance.root", "/path/to/lance/database") \
    .config(
        "spark.sql.catalog.memory_lance.index_cache_backend",
        "moka://?capacity=6442450944",
    ) \
    .config("spark.sql.catalog.other_lance", "org.lance.spark.LanceNamespaceSparkCatalog") \
    .config("spark.sql.catalog.other_lance.impl", "dir") \
    .config("spark.sql.catalog.other_lance.root", "/path/to/lance/database") \
    .config(
        "spark.sql.catalog.other_lance.index_cache_backend",
        "other://?option=value",
    ) \
    .getOrCreate()

For structured Java SDK configuration, build a Session with CacheBackendConfig and register it before the catalog is first used:

import org.lance.CacheBackendConfig;
import org.lance.Session;
import org.lance.spark.LanceRuntime;

Session session = Session.builder()
    .indexCacheBackend(
        CacheBackendConfig.builder("moka")
            .option("capacity", "6442450944")
            .build())
    .metadataCacheBackend(
        CacheBackendConfig.builder("moka")
            .option("capacity", "1073741824")
            .build())
    .build();
LanceRuntime.registerSession("lance", session);

Registration transfers session ownership to LanceRuntime; do not close it directly. In cluster mode, programmatic registration must run in every JVM that accesses Lance. Catalog URI options are recommended when the same configuration should be distributed automatically.

The corresponding environment fallbacks are LANCE_INDEX_CACHE_BACKEND and LANCE_METADATA_CACHE_BACKEND. Per-catalog URIs take precedence over environment backend URIs. Backend URIs take precedence over the legacy LANCE_INDEX_CACHE_SIZE and LANCE_METADATA_CACHE_SIZE settings for the same cache tier.

OpenTelemetry Metrics

Lance Spark can bridge Lance's native Java metrics into the JVM OpenTelemetry pipeline. Enable it with:

--conf spark.lance.otel.enabled=true

You can also set the JVM system property -Dspark.lance.otel.enabled=true, or set LANCE_SPARK_OTEL_ENABLED=true in each JVM environment. When multiple sources are set, Spark configuration takes precedence over the JVM system property, which takes precedence over the environment variable. Values must be true or false (case-insensitive); invalid values disable the bridge and produce a warning.

The bridge uses the OpenTelemetry GlobalOpenTelemetry meter provider, so configure your exporter, resource, and reader with the standard OpenTelemetry Java SDK settings. Spark driver and executor JVMs are independent processes; set the Spark conf or environment variable for every process that should emit Lance metrics.

Example Namespace Implementations

Directory Namespace

import org.apache.spark.sql.SparkSession

val spark = SparkSession.builder()
    .appName("lance-dir-example")
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog")
    .config("spark.sql.catalog.lance.impl", "dir")
    .config("spark.sql.catalog.lance.root", "/path/to/lance/database")
    .getOrCreate()
import org.apache.spark.sql.SparkSession;

SparkSession spark = SparkSession.builder()
    .appName("lance-dir-example")
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog")
    .config("spark.sql.catalog.lance.impl", "dir")
    .config("spark.sql.catalog.lance.root", "/path/to/lance/database")
    .getOrCreate();
spark-shell \
  --packages org.lance:lance-spark-bundle-3.5_2.12:0.4.0 \
  --conf spark.sql.catalog.lance=org.lance.spark.LanceNamespaceSparkCatalog \
  --conf spark.sql.catalog.lance.impl=dir \
  --conf spark.sql.catalog.lance.root=/path/to/lance/database
spark-submit \
  --packages org.lance:lance-spark-bundle-3.5_2.12:0.4.0 \
  --conf spark.sql.catalog.lance=org.lance.spark.LanceNamespaceSparkCatalog \
  --conf spark.sql.catalog.lance.impl=dir \
  --conf spark.sql.catalog.lance.root=/path/to/lance/database \
  your-application.jar

Directory Configuration Parameters

Parameter Required Description
spark.sql.catalog.{name}.root ✗ Storage root location (default: current directory)

Example settings:

from pyspark.sql import SparkSession

spark = SparkSession.builder \
    .appName("lance-dir-local-example") \
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog") \
    .config("spark.sql.catalog.lance.impl", "dir") \
    .config("spark.sql.catalog.lance.root", "/path/to/lance/database") \
    .getOrCreate()
from pyspark.sql import SparkSession

spark = SparkSession.builder \
    .appName("lance-dir-minio-example") \
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog") \
    .config("spark.sql.catalog.lance.impl", "dir") \
    .config("spark.sql.catalog.lance.root", "s3://bucket-name/lance-data") \
    .config("spark.sql.catalog.lance.storage.access_key_id", "abc") \
    .config("spark.sql.catalog.lance.storage.secret_access_key", "def")
    .config("spark.sql.catalog.lance.storage.session_token", "ghi") \
    .getOrCreate()
from pyspark.sql import SparkSession

spark = SparkSession.builder \
    .appName("lance-dir-minio-example") \
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog") \
    .config("spark.sql.catalog.lance.impl", "dir") \
    .config("spark.sql.catalog.lance.root", "s3://bucket-name/lance-data") \
    .config("spark.sql.catalog.lance.storage.endpoint", "http://minio:9000") \
    .config("spark.sql.catalog.lance.storage.aws_allow_http", "true") \
    .config("spark.sql.catalog.lance.storage.access_key_id", "admin") \
    .config("spark.sql.catalog.lance.storage.secret_access_key", "password") \
    .getOrCreate()

REST Namespace

Here we use LanceDB Cloud as an example of the REST namespace:

spark = SparkSession.builder \
    .appName("lance-rest-example") \
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog") \
    .config("spark.sql.catalog.lance.impl", "rest") \
    .config("spark.sql.catalog.lance.headers.x-api-key", "your-api-key") \
    .config("spark.sql.catalog.lance.headers.x-lancedb-database", "your-database") \
    .config("spark.sql.catalog.lance.uri", "https://your-database.us-east-1.api.lancedb.com") \
    .getOrCreate()
val spark = SparkSession.builder()
    .appName("lance-rest-example")
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog")
    .config("spark.sql.catalog.lance.impl", "rest")
    .config("spark.sql.catalog.lance.headers.x-api-key", "your-api-key")
    .config("spark.sql.catalog.lance.headers.x-lancedb-database", "your-database")
    .config("spark.sql.catalog.lance.uri", "https://your-database.us-east-1.api.lancedb.com")
    .getOrCreate()
SparkSession spark = SparkSession.builder()
    .appName("lance-rest-example")
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog")
    .config("spark.sql.catalog.lance.impl", "rest")
    .config("spark.sql.catalog.lance.headers.x-api-key", "your-api-key")
    .config("spark.sql.catalog.lance.headers.x-lancedb-database", "your-database")
    .config("spark.sql.catalog.lance.uri", "https://your-database.us-east-1.api.lancedb.com")
    .getOrCreate();
spark-shell \
  --packages org.lance:lance-spark-bundle-3.5_2.12:0.4.0 \
  --conf spark.sql.catalog.lance=org.lance.spark.LanceNamespaceSparkCatalog \
  --conf spark.sql.catalog.lance.impl=rest \
  --conf spark.sql.catalog.lance.headers.x-api-key=your-api-key \
  --conf spark.sql.catalog.lance.headers.x-lancedb-database=your-database \
  --conf spark.sql.catalog.lance.uri=https://your-database.us-east-1.api.lancedb.com
spark-submit \
  --packages org.lance:lance-spark-bundle-3.5_2.12:0.4.0 \
  --conf spark.sql.catalog.lance=org.lance.spark.LanceNamespaceSparkCatalog \
  --conf spark.sql.catalog.lance.impl=rest \
  --conf spark.sql.catalog.lance.headers.x-api-key=your-api-key \
  --conf spark.sql.catalog.lance.headers.x-lancedb-database=your-database \
  --conf spark.sql.catalog.lance.uri=https://your-database.us-east-1.api.lancedb.com \
  your-application.jar

REST Configuration Parameters

Parameter Required Description
spark.sql.catalog.{name}.uri ✓ REST API endpoint URL (e.g., https://api.lancedb.com)
spark.sql.catalog.{name}.headers.* ✗ HTTP headers for authentication (e.g., headers.x-api-key)

AWS Glue Namespace

AWS Glue is Amazon's managed metastore service that provides a centralized catalog for your data assets.

spark = SparkSession.builder \
    .appName("lance-glue-example") \
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog") \
    .config("spark.sql.catalog.lance.impl", "glue") \
    .config("spark.sql.catalog.lance.region", "us-east-1") \
    .config("spark.sql.catalog.lance.catalog_id", "123456789012") \
    .config("spark.sql.catalog.lance.access_key_id", "your-access-key") \
    .config("spark.sql.catalog.lance.secret_access_key", "your-secret-key") \
    .config("spark.sql.catalog.lance.root", "s3://your-bucket/lance") \
    .getOrCreate()
val spark = SparkSession.builder()
    .appName("lance-glue-example")
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog")
    .config("spark.sql.catalog.lance.impl", "glue")
    .config("spark.sql.catalog.lance.region", "us-east-1")
    .config("spark.sql.catalog.lance.catalog_id", "123456789012")
    .config("spark.sql.catalog.lance.access_key_id", "your-access-key")
    .config("spark.sql.catalog.lance.secret_access_key", "your-secret-key")
    .config("spark.sql.catalog.lance.root", "s3://your-bucket/lance")
    .getOrCreate()
SparkSession spark = SparkSession.builder()
    .appName("lance-glue-example")
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog")
    .config("spark.sql.catalog.lance.impl", "glue")
    .config("spark.sql.catalog.lance.region", "us-east-1")
    .config("spark.sql.catalog.lance.catalog_id", "123456789012")
    .config("spark.sql.catalog.lance.access_key_id", "your-access-key")
    .config("spark.sql.catalog.lance.secret_access_key", "your-secret-key")
    .config("spark.sql.catalog.lance.root", "s3://your-bucket/lance")
    .getOrCreate();

Additional Dependencies

Using the Glue namespace requires additional dependencies beyond the main Lance Spark bundle: - lance-namespace-glue: Lance Glue namespace implementation - AWS Glue related dependencies: The easiest way is to use software.amazon.awssdk:bundle which includes all necessary AWS SDK components, though you can specify individual dependencies if preferred

Example with Spark Shell:

spark-shell \
  --packages org.lance:lance-spark-bundle-3.5_2.12:0.4.0,org.lance:lance-namespace-glue:0.3.0,software.amazon.awssdk:bundle:2.20.0 \
  --conf spark.sql.catalog.lance=org.lance.spark.LanceNamespaceSparkCatalog \
  --conf spark.sql.catalog.lance.impl=glue \
  --conf spark.sql.catalog.lance.root=s3://your-bucket/lance

Glue Configuration Parameters

Parameter Required Description
spark.sql.catalog.{name}.region ✗ AWS region for Glue operations (e.g., us-east-1). If not specified, derives from the default AWS region chain
spark.sql.catalog.{name}.catalog_id ✗ Glue catalog ID, defaults to the AWS account ID of the caller
spark.sql.catalog.{name}.endpoint ✗ Custom Glue service endpoint for connecting to compatible metastores
spark.sql.catalog.{name}.access_key_id ✗ AWS access key ID for static credentials
spark.sql.catalog.{name}.secret_access_key ✗ AWS secret access key for static credentials
spark.sql.catalog.{name}.session_token ✗ AWS session token for temporary credentials
spark.sql.catalog.{name}.root ✗ Storage root location (e.g., s3://bucket/path), defaults to current directory

Apache Hive Namespace

Lance supports both Hive 2.x and Hive 3.x metastores for metadata management.

Hive 3.x Namespace

spark = SparkSession.builder \
    .appName("lance-hive3-example") \
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog") \
    .config("spark.sql.catalog.lance.impl", "hive3") \
    .config("spark.sql.catalog.lance.parent", "hive") \
    .config("spark.sql.catalog.lance.hadoop.hive.metastore.uris", "thrift://metastore:9083") \
    .config("spark.sql.catalog.lance.client.pool-size", "5") \
    .config("spark.sql.catalog.lance.root", "hdfs://namenode:8020/lance") \
    .getOrCreate()
val spark = SparkSession.builder()
    .appName("lance-hive3-example")
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog")
    .config("spark.sql.catalog.lance.impl", "hive3")
    .config("spark.sql.catalog.lance.parent", "hive")
    .config("spark.sql.catalog.lance.hadoop.hive.metastore.uris", "thrift://metastore:9083")
    .config("spark.sql.catalog.lance.client.pool-size", "5")
    .config("spark.sql.catalog.lance.root", "hdfs://namenode:8020/lance")
    .getOrCreate()
SparkSession spark = SparkSession.builder()
    .appName("lance-hive3-example")
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog")
    .config("spark.sql.catalog.lance.impl", "hive3")
    .config("spark.sql.catalog.lance.parent", "hive")
    .config("spark.sql.catalog.lance.hadoop.hive.metastore.uris", "thrift://metastore:9083")
    .config("spark.sql.catalog.lance.client.pool-size", "5")
    .config("spark.sql.catalog.lance.root", "hdfs://namenode:8020/lance")
    .getOrCreate();

Hive 2.x Namespace

spark = SparkSession.builder \
    .appName("lance-hive2-example") \
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog") \
    .config("spark.sql.catalog.lance.impl", "hive2") \
    .config("spark.sql.catalog.lance.hadoop.hive.metastore.uris", "thrift://metastore:9083") \
    .config("spark.sql.catalog.lance.client.pool-size", "3") \
    .config("spark.sql.catalog.lance.root", "hdfs://namenode:8020/lance") \
    .getOrCreate()
val spark = SparkSession.builder()
    .appName("lance-hive2-example")
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog")
    .config("spark.sql.catalog.lance.impl", "hive2")
    .config("spark.sql.catalog.lance.hadoop.hive.metastore.uris", "thrift://metastore:9083")
    .config("spark.sql.catalog.lance.client.pool-size", "3")
    .config("spark.sql.catalog.lance.root", "hdfs://namenode:8020/lance")
    .getOrCreate()
SparkSession spark = SparkSession.builder()
    .appName("lance-hive2-example")
    .config("spark.sql.catalog.lance", "org.lance.spark.LanceNamespaceSparkCatalog")
    .config("spark.sql.catalog.lance.impl", "hive2")
    .config("spark.sql.catalog.lance.hadoop.hive.metastore.uris", "thrift://metastore:9083")
    .config("spark.sql.catalog.lance.client.pool-size", "3")
    .config("spark.sql.catalog.lance.root", "hdfs://namenode:8020/lance")
    .getOrCreate();

Additional Dependencies

Using Hive namespaces requires additional JARs beyond the main Lance Spark bundle: - For Hive 2.x: lance-namespace-hive2 - For Hive 3.x: lance-namespace-hive3

Example with Spark Shell for Hive 3.x:

spark-shell \
  --packages org.lance:lance-spark-bundle-3.5_2.12:0.4.0,org.lance:lance-namespace-hive3:0.3.0 \
  --conf spark.sql.catalog.lance=org.lance.spark.LanceNamespaceSparkCatalog \
  --conf spark.sql.catalog.lance.impl=hive3 \
  --conf spark.sql.catalog.lance.hadoop.hive.metastore.uris=thrift://metastore:9083 \
  --conf spark.sql.catalog.lance.root=hdfs://namenode:8020/lance

Hive Configuration Parameters

Parameter Required Description
spark.sql.catalog.{name}.hadoop.* ✗ Additional Hadoop configuration options, will override the default Hadoop configuration
spark.sql.catalog.{name}.client.pool-size ✗ Connection pool size for metastore clients (default: 3)
spark.sql.catalog.{name}.root ✗ Storage root location for Lance tables (default: current directory)

Note on Namespace Levels

Spark provides at least a 3 level hierarchy of catalog → multi-level namespace → table. Most users treat Spark as a 3 level hierarchy with 1 level namespace.

For Namespaces with Less Than 3 Levels

Some namespace implementations have a flat 2-level hierarchy of root namespace → table. The LanceNamespaceSparkCatalog provides a configuration single_level_ns to enable single-level mode with a virtual "default" namespace.

DirectoryNamespace: By default, uses multi-level namespace mode with manifest-based storage. Tables are stored with hash-prefixed paths for better scalability.

# Default: multi-level namespace mode with manifest-based storage
spark = SparkSession.builder \
    .config("spark.sql.catalog.lance.impl", "dir") \
    .config("spark.sql.catalog.lance.root", "s3://bucket/lance") \
    .getOrCreate()

# Create namespaces explicitly, then create tables
spark.sql("CREATE NAMESPACE lance.mydb")
spark.sql("CREATE TABLE lance.mydb.users (id INT, name STRING)")
# Enable single-level mode for backward compatibility
spark = SparkSession.builder \
    .config("spark.sql.catalog.lance.impl", "dir") \
    .config("spark.sql.catalog.lance.single_level_ns", "true") \
    .getOrCreate()

# Use the virtual "default" namespace (no CREATE NAMESPACE needed)
spark.sql("CREATE TABLE lance.default.users (id INT, name STRING)")

RestNamespace: If ListNamespaces returns an error, single_level_ns=true is automatically enabled.

For Namespaces with More Than 3 Levels

Some namespace implementations like Hive3 support more than 3 levels of hierarchy. For example, Hive3 has a 4 level hierarchy: root metastore → catalog → database → table.

To handle this, the LanceNamespaceSparkCatalog provides parent and parent_delimiter configurations which allow you to specify a parent prefix that gets prepended to all namespace operations.

For example, with Hive3:

  • Setting parent=hive (using default parent_delimiter=.)
  • When Spark requests namespace ["database1"], it gets transformed to ["hive.database1"] for the API call
  • This allows the 4-level Hive3 structure to work within Spark's 3-level model

The parent configuration effectively "anchors" your Spark catalog at a specific level within the deeper namespace hierarchy, making the extra levels transparent to Spark users while maintaining compatibility with the underlying namespace implementation.

Branch Read Option

Set branch to read the current head of a named branch. Do not set version on the same read.

df = spark.read \
    .format("lance") \
    .option("branch", "audit") \
    .load("/path/to/dataset.lance")

Memory Configuration

Lance Spark uses Arrow for data transfer between native code and Spark, and maintains caches for improved performance.

Arrow Allocator

Set via environment variable LANCE_ALLOCATOR_SIZE (default: unlimited).

Controls the maximum memory allocation for Arrow buffers used in data transfer between Lance native code and Spark.

Environment Variable Default Description
LANCE_ALLOCATOR_SIZE Long.MAX_VALUE Arrow allocator size in bytes (global).
export LANCE_ALLOCATOR_SIZE=4294967296  # 4GB

Caching

Lance Spark maintains index and metadata caches to minimize redundant I/O. Cache sizes are configured via environment variables:

Environment Variable Default Description
LANCE_INDEX_CACHE_SIZE 6GB Index cache size in bytes.
LANCE_METADATA_CACHE_SIZE 1GB Metadata cache size in bytes.

For details on how caching works and tuning recommendations, see Performance Tuning - Caching.

Blob v2 Reads

Lance exposes blob v2 columns as struct<kind:short, position:long, size:long, blob_id:long, blob_uri:string>. Descriptor queries do not fetch bytes.

SELECT id, payload.size, payload.kind FROM lance.ns.tbl;

A column is blob v2 when the Arrow field has ARROW:extension:name = lance.blob.v2.

SQL filter pushdown is disabled on blob v2 tables; zonemap fragment pruning still runs. See Blob v2 Writes for copying blobs between tables.

Blob v2 Writes

To write blob v2 columns, set file_format_version to 2.2 or higher and <column>.lance.encoding = blob in TBLPROPERTIES.

Spark accepts BINARY on write. The connector maps that to the Arrow blob write struct at encode time. On read the same columns come back as descriptor structs (Blob v2 Reads).

Copy blobs between Lance tables with a direct column select in INSERT ... SELECT or CTAS. See INSERT INTO and CREATE TABLE.

CREATE TABLE lance.mydb.users (
    id INT NOT NULL,
    content BINARY
) USING lance
TBLPROPERTIES (
    'content.lance.encoding' = 'blob',
    'file_format_version' = '2.2'
);

With file_format_version = '2.2' or higher, blob columns are written using blob v2 encoding and ARROW:extension:name = lance.blob.v2 metadata. The stable and next release selectors resolve to 2.2 or newer, so they select blob v2 as well.

When neither the table nor the catalog sets file_format_version, the table follows Lance's default version, which is 2.2, so its blob columns are blob v2. A catalog-level file_format_version is applied first, so a catalog pinned to 2.0, 2.1 or legacy still produces v1 blob columns.

Creating a table from a schema read back from an existing v1 blob table is the exception: that schema already carries lance-encoding:blob = true, which Lance rejects from 2.2 on, so the new table is created at 2.1 and its blob columns stay v1. Pinning a version of 2.2 or newer for such a schema is rejected — rebuild the column as blob v2 instead. A schema that mixes v1 and v2 blob columns is rejected as well, because no single file format version can store both.

With an older version, such as 2.0, 2.1 or legacy, blob columns use the legacy v1 encoding with lance-encoding:blob = true metadata. Pin one of those versions to keep a blob column on v1 and read it back as BINARY instead of a descriptor struct.

Blob v2 writes must go through the catalog path. Use SQL DDL with TBLPROPERTIES, as shown above, or use the DataFrameWriterV2 API:

df.writeTo("lance.ns.users") \
    .tableProperty("content.lance.encoding", "blob") \
    .tableProperty("file_format_version", "2.2") \
    .create()

Setting only file_format_version does not enable blob encoding. Without <column>.lance.encoding = blob, the column is written as plain BINARY.