toggle

How to Extract SAP HANA Data into a Modern Data Platform

Learn the main ways to ingest data from SAP HANA, including JDBC, ODBC, batch extraction, CDC, native SAP interfaces, and modern ingestion tools. Compare each approach for speed, complexity, data freshness, and on-prem requirements, then see how to choose the right architecture for your analytics or AI environment.

How to Extract SAP HANA Data into a Modern Data Platform

Kannabiran

Sep 15, 2027 |

6 mins

How to Extract SAP HANA Data into a Modern Data Platform

Why SAP HANA Data Ingestion Matters for Modern Analytics

SAP HANA is an in-memory, columnar database at the core of SAP's enterprise application stack — including S/4HANA (the ERP suite), BW/4HANA (the analytics warehouse), and a wide range of custom SAP applications. That distinction matters: HANA is the database engine, not the application itself. When data engineers talk about extracting data from SAP, they're often dealing with HANA as the underlying store, regardless of which SAP product sits on top.

SAP estimates that roughly 77% of the world's transaction revenue touches an SAP system at some point — purchasing, finance, supply chain, and manufacturing included. That data is enormously valuable for analytics and AI, but it largely sits locked inside an ERP core that was never designed with modern lakehouses in mind.

Enterprises building AI models, migrating to lakehouse architectures, or enabling real-time reporting all hit the same wall: they need reliable, governed access to SAP data in platforms like Microsoft Fabric and Databricks. Understanding how to ingest data from SAP HANA database systems into a modern lakehouse is now a foundational capability for data engineering teams. This article covers the connector options available, the ingestion patterns that fit different latency and volume requirements, and how to accelerate delivery without building everything from scratch.

SAP HANA Ingestion Methods: Connector and Approach Comparison

No single extraction method works for every SAP HANA scenario. The right choice depends on how frequently you need data, whether true change capture is required or a full snapshot is acceptable, and how much operational complexity your team can realistically sustain.

Method 

Best for 

Delta/CDC support

Complexity

Typical latency

JDBC/ODBC Driver 

Bulk extracts, ad hoc queries, one-time migrations

None built-in; incremental logic requires custom watermarking (e.g., timestamp columns)

Low–Medium

Minutes–hours

SAP HANA Native Connector (ADF/Fabric) 

Structured ELT from HANA into Azure data stores with schema mapping

Full loads and query-based watermarking; does not use SAP ODP delta queues

Medium

Minutes–hourly

SAP CDC Connector (ODP-based) 

Enterprise-grade delta extraction from ABAP-based SAP systems (ECC, S/4HANA, BW/4HANA)

Native CDC via ODP delta queues; supports full, full-then-incremental, and incremental-only modes

Medium–High

Near-real-time to hourly

Smart Data Integration (SDI) 

Real-time replication into HANA from non-SAP sources (databases, files, apps)

Built-in change capture via remote subscriptions; adapters listen at the source and push changes to HANA

High

Seconds–minutes

OData / REST APIs

Low-volume service integration, self-service apps, fine-grained entity access

No native CDC; client must implement delta logic via query filters or change timestamps

Low–Medium

Per-call (near real-time); overall refresh tied to client schedule

Third-party replication tools (Qlik, Fivetran, etc.) 

Turnkey ingestion into cloud warehouses with managed monitoring

Vendor-specific log-based or trigger-based CDC; specifics vary

Medium

Near-real-time to hourly

A few trade-offs worth internalizing before you pick:

Full load vs. incremental/CDC. JDBC/ODBC and native HANA connectors are straightforward for full snapshots. They become impractical as tables grow into hundreds of millions of rows and refresh windows shrink. The SAP CDC connector sidesteps this by subscribing to ODP delta queues — SAP maintains the change log, and your pipeline retrieves only what changed since the last run. Third-party tools take a similar approach using transaction logs or triggers. For teams that need SAP HANA data replication at scale, ODP-based CDC connectors and third-party replication tools are typically the most operationally sustainable path.

On-premise vs. cloud HANA. JDBC/ODBC and OData work the same way regardless of where HANA is deployed, as long as network connectivity is in place. The SAP CDC connector supports ODP-enabled SAP systems whether they run on-premises, on Azure, or via RISE with SAP.

SDI is directionally different. It's primarily designed to replicate data into HANA from external sources — not out of it. If the goal is extracting from HANA into a lakehouse, SDI isn't pointed in the right direction.

Two scenarios that make the choice concrete:

  • A finance team needing hourly deltas of General Ledger data from S/4HANA into Azure would reach for the SAP CDC connector. GL data is exposed via ODP-enabled ABAP CDS views; the connector subscribes, takes an initial full snapshot, then pulls only new and changed records on each subsequent run — no custom merge logic required.

  • A team running a one-time historical migration — say, five years of transactional records — would reach for JDBC/ODBC. No change tracking needed, no ODP configuration required. Run batch selects partitioned by fiscal period, load the target, and you're done.

The method comparison above gives you a solid decision anchor. The next section walks through the implementation steps for whichever path fits your scenario.

Step-by-Step: How to Ingest Data from SAP HANA

The steps below are deliberately platform-agnostic — the same sequence applies whether you're wiring up Azure Data Factory, Microsoft Fabric, or Databricks. The connectors and UI differ across platforms, but the underlying building blocks — network access, drivers, credentials, extraction logic, and monitoring — stay consistent.

1. Validate prerequisites: HANA version, user permissions, and network access

Start by confirming which version of HANA you're connecting to — HANA 2.0 (on-premise) or HANA Cloud — because the client driver requirements differ between them. Create a dedicated technical user with SELECT on the target tables or views, plus access to catalog metadata. On the network side, verify that your compute layer can reach the HANA indexserver SQL port in the 3xx15 range — for example, port 30015 on a single-container system.

2. Choose your connector based on source type and latency needs

For batch extraction of tables and calculation views, an ODBC/SQL connector using the HANA client is the straightforward choice. If you need SAP HANA real-time data ingestion from SAP application sources — ECC, S/4HANA, BW — you'll want an ODP-based CDC connector backed by the SAP .NET Connector.

Don't over-engineer this decision. A full-load migration doesn't need CDC infrastructure, and a finance team running hourly delta loads of GL data shouldn't be polling with a raw JDBC query.

3. Deploy a Self-Hosted Integration Runtime for on-premise HANA

For on-premise HANA, you need a runtime agent that bridges your private network and your cloud platform. Deploy a Self-Hosted Integration Runtime (SHIR) — or an equivalent self-hosted gateway — on a Windows Server close to your HANA host.

A practical setup: provision a Windows Server VM in Azure, connect it to your on-premise network via VPN or ExpressRoute in a peered VNet, and install the SHIR service there. Size it at a minimum of 4–8 vCPUs and 8–16 GB RAM, and ensure outbound HTTPS (port 443) to your cloud platform and inbound-to-HANA access on the appropriate 3xx15 port.

4. Install SAP HANA client drivers or SAP .NET Connector

On the SHIR host — or wherever your compute runs — install the SAP HANA Client 2.0, which includes ODBC drivers for both Windows and Linux. For CDC pipelines using ODP, you'll also need the SAP Connector for Microsoft .NET (NCo 3.0) with its assemblies registered in the Global Assembly Cache (GAC) — a step that's easy to miss and produces cryptic errors when skipped.

Match the driver bitness (64-bit) to your runtime environment, and verify the installation with a quick HANA client connection test before moving forward.

5. Configure the linked service or connection with secure credentials

Define your connection object — a linked service in ADF/Fabric, a DSN or connection string in Databricks — using the HANA ODBC driver, hostname, instance number, and port. Always store credentials in a dedicated secret store: Azure Key Vault, Databricks Secrets, or Fabric's equivalent.

Hard-coded passwords in pipeline JSON are an audit finding and a support ticket waiting to happen. For CDC connections, you'll also configure the ODP subscriber name and endpoint here, following your platform's vendor documentation.

6. Build and parameterize the extraction pipeline

Design your pipeline around parameterized SQL or named views rather than SELECT *. Apply server-side filters — date ranges, key ranges, relevant columns — to push as much work as possible to HANA's columnar engine.

For large tables, build partitioned reads that split extraction by a date or numeric key column and run parallel threads, rather than issuing a single long-running query. For incremental loads, implement a high-water mark pattern — tracking the last extracted timestamp or document number — or use ODP subscriptions so subsequent runs pull only changed records.

7. Schedule, monitor, and handle failures

Configure scheduled triggers for batch pipelines and event-based triggers for CDC flows, with appropriate retry policies attached. Monitor three things consistently:

SHIR node health

Driver-level errors (especially ODBC timeout and authentication failures)

ODQ subscriber state if you're running CDC

Build failure handling in from the start — exponential backoff for transient network issues, dead-letter storage for malformed records, and idempotent re-run logic so a failed execution can be safely restarted without duplicating data downstream.

Performance, Security, and Common Pitfalls to Avoid

Getting an SAP HANA pipeline to run is a different problem from getting it to run safely at scale. HANA is an in-memory column-store sitting at the heart of your production ERP — a poorly designed extraction can lock rows, spike memory, and degrade live transactional workloads. The tips below are the kind of things that typically get learned the hard way.

How to Optimize SAP HANA Extraction Performance

Partition large extractions, always. A single full-table scan against a 500M-row column-store table is a reliable path to timeouts and out-of-memory errors. Microsoft's SAP HANA connector supports physical partitions, dynamic range, and partition column options that let the engine run multiple queries in parallel rather than one giant sequential pull. Partition on a date or integer column that distributes rows evenly — monthly buckets work well in practice.

Push filtering down to HANA, not up to your runtime. Use WHERE clauses and column selection in your extraction query to reduce the result set before it hits the network. Filtering downstream means you've already paid the cost of moving data you didn't need.

Tune parallelism conservatively. Start with a degree of copy parallelism around 4 and step up gradually while watching HANA memory and CPU metrics. Aggressive parallelism can starve concurrent OLTP workloads or trigger memory pressure — the opposite of what you're after.

Schedule off-peak and use read-optimized views. Long-running extractions on live tables can cause row-level or table-level locks that ripple into production transactions. Where possible, extract from HANA calculation views or reporting views rather than base tables, and schedule heavy loads outside business hours.

Plan for schema drift. HANA schemas evolve — columns get added, data types change, tables get renamed. Build your pipelines to extract against defined views or explicit column lists, and add a compatibility check step so schema changes surface as pipeline failures rather than silent data corruption downstream.

How to Secure SAP HANA Pipelines and Manage Licensing Risk

Use a dedicated, least-privilege technical user. Create a HANA database user specifically for extraction, scoped to read-only access on the schemas and views it needs. Avoid granting broad SYSTEM or DATA ADMIN privileges. Narrow permissions limit the blast radius of a misconfiguration or a compromised secret.

Vault your credentials. Pipelines should retrieve HANA credentials from a secret store — like Azure Key Vault — at runtime, not from hardcoded connection strings or config files checked into source control. Rotate secrets on a defined schedule.

Enforce TLS end-to-end. Configure your HANA connection to require TLS, validate the server certificate, and disable weak cipher suites. When moving large volumes of financial or operational data across network boundaries, encryption in transit isn't optional.

Treat SAP indirect access risk as a legal question. Reading data from SAP HANA into external platforms — a data lake, a custom analytics app, a non-SAP BI tool — can trigger SAP indirect access licensing obligations, depending on how that data is used. The technical architecture alone won't resolve this. Before scaling extractions that feed non-SAP transactional systems, get SAP licensing experts or legal counsel involved. This is a known cost exposure that data engineering teams regularly discover too late.

How datakulture Accelerates SAP HANA Ingestion with Bizweave

Getting data out of SAP HANA cleanly — with delta logic, proper credentials, schema handling, and downstream orchestration — takes more engineering effort than most teams budget for. The connector works; the pipeline around it rarely does on the first try.

datakulture built Bizweave specifically to close that gap: not to replace the native connectors in Fabric or Databricks, but to remove the repetitive, error-prone work that surrounds them. Bizweave functions as a data foundation accelerator — combining data ingestion, data quality, data lineage, orchestration, DataOps, real-time data support, and AI-assisted data engineering into a unified delivery model built specifically for Microsoft Fabric and Databricks.

How Pre-Built SAP Ingestion Patterns Speed Up Delivery

Bizweave ships with pre-built ingestion patterns for common SAP sources, covering full loads, delta extractions, hierarchical structures, and combined application-layer and database-layer scenarios. These patterns aren't just connector templates — each one ties directly into Bizweave's orchestration layer, built-in data quality controls, and lineage tracking, making them reusable production assets rather than starting-point scripts.

Instead of designing pipeline logic from scratch for every SAP object, your team starts from a proven template and configures what to ingest rather than figuring out how. In practice, this has compressed initial pipeline delivery from months to a matter of weeks — particularly in landscapes with multiple SAP modules and heterogeneous schemas.

How AI-Assisted Mapping Reduces Manual Transformation Work

Mapping ABAP-based data structures to modern analytics models is where SAP HANA ingestion projects quietly lose weeks. Bizweave addresses this with AI-assisted mapping that reads source metadata and suggests joins, data types, and transformation logic automatically.

This reduces the manual field-by-field reconciliation that typically follows a successful extraction — and lowers the risk of misreading SAP business semantics when landing data in Fabric or Databricks.

How End-to-End Lineage and Data Quality Support SAP Modernization

One of the least-solved problems in SAP modernization is visibility: knowing exactly how a number in a Fabric report traces back to its originating SAP HANA table. Bizweave captures source-to-target lineage and validates data at each pipeline stage, so teams can detect anomalies before they reach downstream consumers.

This matters especially when decommissioning legacy SAP reports — and when you need to prove equivalence in the new stack before pulling the plug on the old one.

How Bizweave Orchestrates Across Fabric and Databricks

SAP HANA ingestion doesn't end at extraction. It needs to coordinate with transformation jobs, medallion layers, and reporting refreshes running across different platforms. Bizweave provides orchestration that spans SAP extraction through to Microsoft Fabric OneLake and Databricks, handling scheduling, dependency management, and error recovery in a single operational view.

This orchestration aligns with the modern data engineering stacks that surround Fabric and Databricks — including Spark and Delta Lake for processing, Unity Catalog and Lakeflow for governance, and tooling like dbt and Airflow for transformation workflow management. That replaces the stitched-together approach of managing separate tools for each environment.

How Bizweave Embeds SAP Integration Expertise

Bizweave embeds hard-won integration knowledge directly into its accelerators: best-practice patterns for delta handling, historical load safeguards, and reconciling SAP application logic with lakehouse schemas. For datakulture engagements, this means the team focuses on business outcomes — harmonized finance data, reliable manufacturing analytics, AI-ready pipelines — rather than re-solving the same SAP plumbing problems project after project.

Bizweave is designed to turn bespoke SAP ingestion work into a repeatable, accelerator-led implementation model — one that brings DataOps discipline and real-time data readiness to every engagement from the start.

Together, these capabilities make Bizweave a practical accelerator for enterprises that want production-grade SAP HANA pipelines without carrying the full engineering cost of building and maintaining them in-house.

Conclusion: Turning SAP HANA Data into an AI-Ready Foundation

Picking the right connector is the starting point, not the finish line. The engineers and teams who extract sustained value from SAP HANA are the ones who build for scale, governance, and long-term modernization — replacing brittle, one-off exports with governed pipelines that preserve lineage, handle schema drift, and keep data fresh enough to power reliable AI agents. That shift takes more than tooling; it takes deliberate architecture.

That's where datakulture comes in. Through Bizweave, datakulture helps data engineering teams move from fragile SAP pipelines to an AI-ready data foundation — without locking you into a single platform or forcing a rebuild from scratch. If you're ready to turn your SAP HANA data into something your analytics and AI workloads can actually depend on, talk to datakulture's data engineering team.

Frequently asked questions

1. What is the best way to ingest data from SAP HANA into Microsoft Fabric or Databricks?

For Microsoft Fabric, the SAP HANA Database connector in Data Factory pipelines or Dataflow Gen2 is the most straightforward path — it lands data directly into a Lakehouse or Warehouse. For Databricks, the common approach is either JDBC/ODBC connectivity from a Databricks cluster or using an intermediary like Azure Data Factory to move data into cloud storage, then processing it in Databricks with Delta Lake. The right choice depends on your latency requirements, SAP source type (raw HANA tables vs. ERP/BW objects), and whether you need delta extraction.

2. Can I extract data from SAP HANA without impacting production performance?

Yes, but it requires deliberate design. Use a read-only technical user with scoped permissions, partition large table extractions by date or key range to avoid long-running single queries, and schedule heavy loads outside business hours. Where available, read from a HANA secondary or analytics replica rather than the primary production node — that single decision eliminates most contention risk.

3. What is the difference between the SAP HANA connector and the SAP CDC connector?

The SAP HANA connector provides direct database-level access to HANA tables and views via ODBC/JDBC — well-suited for full extracts and simple timestamp-based incremental loads, but with no awareness of SAP business deltas. The SAP CDC connector uses SAP's Operational Data Provisioning (ODP) framework to extract only changed records from SAP ERP and BW sources after an initial full load. Short version: HANA connector for database-level extraction; SAP CDC connector for business-aware, ongoing SAP HANA data replication.

4. Do I need a Self-Hosted Integration Runtime for cloud-hosted SAP HANA?

Usually yes — if your HANA instance sits behind a private network, VPN, or firewall, the SHIR acts as the secure bridge and also hosts the required SAP client drivers. If the HANA instance is internet-accessible (for example, RISE with SAP with appropriate network exposure), native cloud connectors may reach it without an additional runtime. When in doubt, deploying a SHIR on an Azure VM in a peered VNet is the safer, more reliable default.

5. Does extracting data from SAP HANA trigger SAP indirect access licensing?

Not automatically — but the answer depends on how the extracted data is used, not simply that extraction happened. Pulling HANA data into analytics platforms for reporting and analysis is generally treated differently from scenarios where non-SAP systems use SAP data to drive transactional document creation. SAP's Digital Access model addresses indirect document creation specifically. Because licensing terms are contractual and landscape-specific, this is a question to resolve with your SAP account team or licensing advisor — not just your data engineering team.

6. How do I handle incremental loads and change data capture from SAP HANA?

You have two main options depending on your source. For ERP and BW data, use SAP CDC connectors built on the ODP framework — these handle initial full loads and pull only changed records on subsequent runs, making them the preferred approach for SAP HANA real-time data ingestion requirements. For table-level CDC directly in HANA, SAP Smart Data Integration (SDI) can capture insert, update, and delete events via triggers. For simpler scenarios where tables have a reliable timestamp or sequence column, a watermark-based incremental query (WHERE LAST_CHANGED_AT > ?) with a stored high-water mark is often sufficient — and considerably easier to maintain.

7. How long does it typically take to build a production-grade SAP HANA ingestion pipeline?

A proof-of-concept covering connectivity and a handful of tables usually takes one to three weeks. A production-grade pipeline — with proper error handling, monitoring, incremental logic, data quality checks, and governance — typically runs four to twelve weeks for a mid-size SAP landscape. Large enterprises with multiple SAP modules, CDC requirements, and compliance obligations often invest several months to reach a fully hardened, observable pipeline. The single biggest variable is usually schema complexity and the number of SAP source objects, not the connector setup itself.