22.02.2023
The ELT process basically consists of 3 steps:
- Extract – extracting data from sources and supplying a temporary stop.
- Load – uploading data to the destination.
- Transform – a step in which the data is transformed, including their reconciliation, cleaning, and correction.
Traditionally, large organizations that had significant transactions took advantage of ETL (extract, transform, load) to process data in their systems for analysis and reporting purposes. Loading your data into a cloud data warehouse and data lake provides scalability, ease of access, low storage costs, and operational efficiency. With the ability to store data in the cloud and process it, this approach is slowly giving way to data processing after it has been downloaded and replicated to the cloud. Cloud providers even charge separate fees for data storage and processing, giving customers greater flexibility. This is why many users are moving towards the ELT ecosystem.
Why ELTs?
While there are many benefits to implementing ELT, we believe the following three provide organizations with the most value:
- Extract any data from any source at scale and at high speed
Larger enterprises typically have many different data sources such as applications, databases, files, streaming, etc. Using ELT means you can retrieve and replicate data from different datasets, regardless of source, whether structured or unstructured, linked or unbound. - Faster data transformation using cloud computing
The ELT does not have to wait for the data to be transformed and then loaded. The transformation process takes place where the data resides, so it can be accessed in seconds. This is a huge benefit, especially when there is a need to process data in a short time. - Save time and money
ELT reduces data transfer time. It also does not require a temporary data system or additional remote resources to transform data outside the cloud. There is no need to move data to and from cloud ecosystems for analytics, meaning zero cost of data output. It also lowers TCO (Total Cost of Ownership) due to improved efficiency.
Like a platform that manages company data Informatica helps to optimize the ELT?
Smart Cloud for Data Management (IDMC) Informatica is a comprehensive data management platform based on artificial intelligence that offers key functionalities required to optimize ELT processes. In particular, the bulk ingest and upload functions help you to efficiently perform the next steps of the process: extracting, loading and transforming data. Let's take a look at these opportunities and their business value.
Step 1 & 2: Unpack and load
During the ELT extraction step, data is first extracted from one or more sources. This can be IoT data, data from social media platforms, from the cloud or on-premise systems. Then, in the loading step, this data is transported to a data lake or data warehouse. The extraction and loading steps can be efficiently performed by the service Cloud Mass Ingestion (CMI).
CMI can retrieve and replicate unstructured, semi- and fully structured data at scale from a variety of databases, applications, files and data sources, streamed with very low latency, for cloud and messaging purposes. It provides a code-free, wizard-based approach to data ingestion, replication, and data synchronization. CMI also enables both technical and non-technical users to create data pipelines in minutes. Featuring a unified user interface, CMI provides out-of-the-box connectivity to hundreds of sources and destinations.
The highly scalable service can be used to ingest terabytes of almost any data, with almost any pattern and latency. It can do this both in real time and in batch. Because this bulk data ingestion service is part of the broader IDMC platform, it includes native user management, monitoring capabilities, and alerting mechanisms.
Step 3: Transform
During the ELT transformation step, the data is converted from the source format to the format required for further analysis business to take further action. Advanced pushdown optimization (APDO) which is a function of the service Cloud Data Integration Informaticacan help with this transformation. Pushdown optimization is a performance tuning technique. The transformation logic is converted to SQL and passed to the source or destination database, or both.
APDO allows two types of pushdown optimization:
- Data Warehouse Upload uses SQL queries to move data from the staging area to the operational data warehouse (ODS) and from the ODS to the enterprise data warehouse (EDW) within the data warehouse.
- Ecosystem upload uploads data from the cloud data lake to the data warehouse using native ecosystem commands.
Benefits of using APDO have:
- zero data outbound costs as data is not moved from the underlying cloud infrastructure,
- faster than traditional ETL,
- independent of the ecosystem, which makes it easy to change data warehouse provider,
- easy switching between runtime options,
- extensive support for connectors for all major cloud ecosystems,
- no coding experience is required.
How can the combination of CMI and APDO optimize ELT processes?
1. When you store data directly in a cloud data warehouse
Many organizations store data from many different local and cloud sources in a cloud data warehouse. However, before this data is used for business analysis, it is transformed into a data warehouse.
In this scenario, CMI Informatica can be used to ingest or replicate data from various streaming sources, applications, or relational database sources, to the staging area of a cloud data warehouse such as Snowflake, Google Big Query, Amazon Redshift, Azure Synapse, or Databricks. APDO can then be used to transform this staging data into a data warehouse through a data warehouse pushdown.
With this approach, data can be delivered to the data warehouse from multiple endpoints at high speed. By leveraging existing computing power, this maximizes the value of your existing cloud data warehouse investments. This also eliminates any additional data transfer costs.
Illustration of storing data directly in a data warehouse in the cloud
2. When you store data in the cloud data lake before moving it to the cloud data warehouse
Many organizations choose to first store data from many different on-premise and cloud sources in a cloud data lake because, unlike a data warehouse, it provides them with lower-cost, large-scale storage and the flexibility to store unstructured and semi-structured (hierarchical) data. This data is then transformed before being stored in the data warehouse.
In this scenario HCM Informatica can be used to ingest or replicate data from various streaming sources, applications, or relational databases to a cloud data lake such as Amazon S3, Azure Data Lake Storage, or Google Cloud Storage. Before replicating the data to the data warehouse using the pushdown ecosystem, APDO can be used to transform this data in the cloud ecosystem.
Illustration of storing data in a data lake in the cloud before moving it to a data warehouse in the cloud
With this approach, data is delivered to the data lake from multiple endpoints at high speed. In this case, data transfer is free. There is also better performance, resulting in fewer computing hours. This means significant cost savings.
ELT optimization with CMI and APDO
CMI can deliver high-speed data from a wide variety of data sources with minimal transformation. APDO, on the other hand, can process data faster with zero egress charges. The combination of these two IDMC services will optimize the ELT process, saving time and reducing TCO. Learn more about HCM i APDO just now.
Read more here.




