Qlik Open Lakehouse Deep Dive
In this session at Tech Field Day Experience at Qlik Connect 2025, Qlik unveiled and explored its new Qlik Open Lakehouse initiative, an outcome of the company’s acquisition of Upsolver. The presentation focused on the growing market adoption of Apache Iceberg, an open table format for data lakes, driven by its advantages in cost savings and interoperability. Ori Rafael highlighted the transition from traditional, tightly coupled data warehouses to decoupled data lakehouses using Iceberg. This shift enables enterprises to eliminate redundant warehouse storage costs by directly writing to cost-efficient object storage like Amazon S3. Furthermore, Iceberg allows customers to decouple compute from storage and use multiple analytics engines (like Snowflake for BI or Databricks for AI) on a unified data layer, enabling a truly open and flexible data architecture. Qlik’s acquisition of Upsolver has enhanced its capabilities to deliver an enterprise-grade, high-performance lakehouse solution. The Qlik Open Lakehouse, now integrated into Qlik’s Talend Cloud platform, provides high-throughput ingestion, real-time processing, and automatic optimization of Iceberg tables. This includes features such as adaptive file compaction, dynamic partitioning, and efficient snapshot cleanup to keep storage lean and query performance high. Ori shared benchmarks demonstrating that Iceberg tables managed by Upsolver could achieve query performance nearly on par with native Snowflake storage, while offering as much as 2x improvement in storage efficiency over other Iceberg implementations. This level of performance addresses previous shortcomings of Hadoop-based data lakes and provides a practical, streamlined experience that doesn’t require specialized big data engineering expertise. Antoine Richard and Vijay Raja elaborated on Qlik Open Lakehouse’s integration into Qlik Talend Cloud, detailing its core capabilities that support ingestion from hundreds of sources via batch or CDC, optimizer services tailored for Iceberg, and seamless mirroring into Snowflake without duplicating data. They outlined support for Iceberg catalogs like AWS Glue, Polaris, and Snowflake Open Catalog, enhancing query engine compatibility across platforms like Spark, Trino, and Dremio. The product roadmap includes enhancements like streaming ingestion, support for additional cloud providers, and advanced transformation tooling, reinforcing Qlik’s mission to provide an end-to-end data integration platform. A recorded demo concluded the session, illustrating how users can build data pipelines into Iceberg via Qlik Talend Cloud, view data in Athena or Snowflake, and perform seamless transformations, all while maintaining a familiar UI and minimizing the need for warehouse compute resources.
Presented by Ori Rafael, Senior Director of R&D, Vijay Raja, Product Marketing Director, and Antoine Richard, Senior Principal Product Manager, Qlik. Recorded live in Orlando, Florida on May 12, 2025 as part of Qlik Connect 2025. Watch the entire presentation at https://techfieldday.com/event/qlikconnect25/ or visit https://TechFieldDay.com for more information.
Transcript
My name is, uh, I was the CEO and co-founder of Olver, was recently acquired by Qlik. Today I lead the UPS solver group within Qlik. Today, we came to talk to you about the Lakehouse, the iceberg Lakehouse.
I see that you have a lot of questions, which is great. Uh, always the best way to convey material like that. And I wanna start by talking about the demand for iceberg.
Like we've been talking to customers, building data lakes and data warehouses for quite a few years, and we are definitely seeing a wave, not a small wave happening with the iceberg. If, uh, we look at the history of data warehouses, we used to have the transition to the cloud where everyone were running to buy something like an Amazon Redshift instead of managing their own instance. And then everyone went to a decoupled data warehouse and forgot about the old, manage the cloud data warehouse.
And then they're moving to a decoupled storage layer, which is basically what we are doing with Iceberg. And it's not a small transition. The reason I'm mentioning all of them took together is that we are seeing, uh, uh, everyone run into iceberg.
And that statistics kind of backs this up. This is actually data from last year, uh, from a survey that Remu did among enterprises and among enterprises, 69% think that more than 50% of their analytics will run on iceberg within three years. That's quite a change.
I think maybe even a little aggressive migrations always take longer than you plan, but definitely that's where they're aiming for. So why, why, what's the reason? Why are everyone rushing to iceberg?
Like it's the new gold? And I think that there are eventually two key reasons. One is cost, and the other is no locking or interoperability, depending how you wanna define it.
With cost. Every time that you use a data warehouse and with a data integration company, know that you write into a data warehouse, you pay money for the data warehouse, there is a cluster running that spins up. By the way, in as a c as a company that's doing a lot of CDC, we often see those costs really pile up because you want to get your data in real time and the cost of merges into data warehouses or something that's very expensive.
And now there is a, a paradigm shift. Instead of paying the data warehouse when you're ingesting, I can just write to S3. I don't need to pay to anyone.
I can take an entire piece of cost that's very, that's very, uh, material in the data warehouse bill and just eliminate it. And it's very easy for someone that owns the data warehouse in, in, in an enterprise to understand that it's not, Hey, I'm going to reduce costs. Some kind of amorphic promise that no people don't understand.
I'm shutting this down. That cluster does not exist anymore. Very easy to, to understand the cost, uh, the cost benefit.
And that's without going into philosophy of there is more, more, uh, let's say negotiation power with the data warehouse vendor. Just talking about the, the cost that you shut down. And that takes us to think to the second reason, and I don't think it's less important than the first reason.
And the interoperability. And the reason that's so important for everyone right now is ai. Maybe I would, uh, meet you a few years ago and I would say, well, you're not gonna have just one data platform for all your data.
And people would argue no one-stop shop is great. People forgot the lesson from their Oracle days that it's very hard to be tied into one technology. I was an Oracle DBAI remember using a lot of Oracle tech that I don't necessarily saw as best of breed because my data was there.
You, you are first gonna try a tool that's gonna be close to where your data gravity is, where your storage is, and iceberg basically unlocks it from that. So if tomorrow I want to use Databricks for AI and Snowflake for bi, which is something that we see all the time, I can do that with same copy of data. I can choose best of breed without the burden of my legacy and being tied to a specific format.
I think those together are the two main reasons that are pushing people to iceberg. And kinda, if you're not convinced, it's easy to just look at what happened in the last year or so. Every day, big data platform vendor has made a big play around Iceberg.
And if you're gonna listen to their earnings call, they're gonna talk about Iceberg as one of the, uh, uh, key key initiative. So that's including Snowflake, but that includes also Databricks that made a big acquisition and went into that market and didn't want to let the Iceberg Revolution pass them. Although they had Data Lake, and you can see it in Amazon, that, uh, releasing Iceberg products, um, left and right, like every one of the major data platform companies or cloud providers as standardizing on Iceberg as their first, uh, priority format.
And you can say that for some, maybe the second, but for most it's just the first. So the entire market accepted this as the standard of how we are going to build the lacoss, which is great. I founded Olver many years ago.
I've been waiting for maybe four or five years for something like this to happen. But it's great that it's finally here. Alright, so what is actually a lakehouse?
I think the easiest way to understand a Lakehouse is it's just a decoupled database. You take all of those things that you see here on the, on the screen used to be just one box in my organization, and now we're separating all of them into separate pieces. The piece I think that's, uh, harder, hardest for customers to actually get right is storage.
You can say that's the underneath the storage layer. You could include the files that you store, the metadata that you store using Iceberg, and maybe even the, the, the catalog itself. All of these are now can I can buy this from here and this from here.
And it's all very confusing from the customer. And the second thing, it's also harder because the, the IP of, uh, optimizing a file system for good performance is something that database companies have been doing for quite a while. And that's really their core ip.
And now you're gonna say, Mr. And Mrs. Customer, you're now responsible.
And if you're not gonna store your data in in the right way, you're not gonna get the same performance you you got in your data warehouse. And that's a recipe for failing. And why am I saying it's a recipe for failing?
Because we already tried it. There was a dupe, a dupe started with the promise of replacing the data warehouse. Why didn't it fail?
Why, why did it fail? Because it sucked. It was slow and hard to to use.
Like eventually users want the same thing that they had with a data warehouse. They want to the, the cost advantage of a, of the, the lake, but they want the ease of use and the performance of the warehouse. That's what's actually making it into, into a lakehouse.
Uh, so now, how will a customer do this? All of this. And the second question, which vendors are going to help?
Because in the past, no one can really touch what was most precious for the data warehouse company. But now, if I'm a customer, I'm thinking, well, if my problem is that I'm going to be locked to a specific data warehouse, should I let the data warehouse now govern all of this open source thing? Will won't.
I find myself in the same place I've been, like the, I'm doing a project for the same thing I'm trying to solve. So what will I actually do? I can actually go to a data integration company.
That leaves off the idea that there are gonna be multiple data platform that has every incentive to create a data layer that all the different data platforms could be, could be reading from. So now it's not only a question of how I build a warehouse in a different way, but who am I going to build this with? And now I also have the option, we have to talk about it.
The customer can decide, I can do this myself. For some customers, that's fine. That's, uh, they, they welcome the engineering, uh, complexity.
Alright, so just repeating the, a little bit of the complexity. Optimizing and, and tuning file systems. Very hard work.
I I've been doing this for quite a few years. Very hard work to do. Uh, cleanup, very, something that sounds very intuitive.
People may remember this from the days that you used to run vacuums on data warehouses and data, uh, engineering teams used to complain a lot about those maintenance tasks that they need to do. So think about it, that every time I'm gonna go to Iceberg and I'm going to perform a commit operation, I'm going to create an old snapshot. I'm going to create orphan files.
All of this is going to create a lot of garbage. That garbage is gonna slow down my performance and it's gonna cost me more in storage. So I need to clean all of that in a, in a, in a good way.
I need to suddenly integrate with multiple catalogs because I want to run the data in one place and I want Snowflake and Databricks and Amazon and like everyone to be able to read it. But there are differences between the catalog. So I can, for example, write data in a way that Amazon can read it and Snowflake will not be able to read it.
True story that influenced the way we are going to releases on the click, uh, on the click cloud. Uh, the last one that I want to talk about maybe in a second, I know we don't have a lot of time merge operations. That's one of the big ones we see it a lot is, is a company that's doing, uh, CDC merge operations.
Were very expensive on the data warehouse on Iceberg. I would say we are still in the early days of doing merge operations in real time and in high scale. And you need to engineer your way around it if that's what you wanna do.
That's something that, uh, we've been spending quite a few months on at Qlik. Alright, so January, Qlik acquires Olver in order to go, uh, deeper into the, uh, into the Lakehouse. Uh, so I wanted to maybe give you some visibility why Qlik decided to do it.
And I think maybe a a a, a good start was experience. So Olver has been building what you can compare to Iceberg. So data lake tables.
Olver was the, uh, combination of ETL with data lakes. I wanna stream high scale data into a data lake and query it in the same level of performance I had in my data warehouse since 2017. We were kind of beating our head around, how are you going to build a file system that's going to perform well over object storage?
Uh, I think that's before Iceberg even even started when Iceberg came out. It was the standard we've been waiting for and we of of course adopted all our product into for Iceberg. Uh, the second reason was high throughput.
So Qlik is a company that works with big data, but VER was a company that worked, worked with really big data. So if Qlik would, would would work at the scale of 10, tens of thousands of events per second. Olver worked with millions of events per second and added new sources into the Qlik ecosystem.
For example, streaming and object stores with data at very high scale that would be processed in the cloud, in the cloud native way. So you can basically scale with the cloud for, uh, as many events per second or whatever workload you wanna, uh, you wanna, we wanna support. So that's something that was important for Qlik.
So one story was the Lakehouse. How is Qlik going to come into the Lakehouse and have a best of breed, uh, best of breed product. Product.
And the second, how are we going to process very big data and support streaming workloads, which is WhatsApp app solver did. Alright, so with this slide, I'll try to sum up the offering that we used to have on app solver product, not the click product yet. We're gonna talk about the click product a little later.
But on the app solver product, you can divide it into three sub-services that were used and are going to be reintroduced into the click platform. The sec, the first is ingestion. I'm going to choose a source and now I want that source to be an iceberg table in my lake.
I don't wanna start spending a lot of time on scaling and engineering and like all the nitty gritty stuff, we call that zero ETL because it was really zero tl all you defined is your source and your target and all the rest happened automatically including schema evolution. So the first, we have always have to start with ingest. The second is once that I actually have a table, Olver had an entire part, entire component just for managing iceberg tables.
By the way. You didn't have an, you didn't have to write those tables with Olver in order to manage them with Olver. So it only that component of olver connected to the Iceberg catalog and to S3 looked at the file system and there were a bunch of parameters that influenced how it work.
How many files, what are the size of the files and what velocity are they writing. Like there is, it's just, uh, you can call it an equation. We've been balancing since the, uh, 2017.
And I'll share some results in, uh, in a couple of minutes benchmark, uh, to other, uh, to other solutions. So ingest, optimize, and the last part was actually doing some transformation, but those transformation are not ex exactly the same, like click transformation. So in Qlik in QTC, you're able to write a transformation that will be pushed down into your data warehouse.
With the UPS iceberg live tables. You basically write SQL that's being processed on the app solver engine. So the app solver engine is a streaming engine.
So I can do my data prep with streaming and then I could be using olver and after I'm gonna finish, I can use push down in the same way. So you could compare this to products like Snowflake Dynamic Tables or Databricks live tables. It's a very similar offering.
We've been doing it since 2018, so it's not new to us. The idea, I wanna declare more tables in my medallion architecture, bronze, silver, gold, using SQL and not worry about orchestration and engineering. I just want everything to be declarative.
The declarative is the key, is the key word. And of course there's the catalog and whatever query engine you wanna put on top. Alright, so I, I promise to share some statistics.
So there are two types of statistics on the screen. On the left, you see storage. So I used to have a table, uh, that was 10, uh, 10 terabyte, and now I'm gonna move to Iceberg.
And suddenly that table is gonna be 50 terabytes not cool. Like, that's not how I wanna work. And the closer you get to real time, the more you have commits, the more you're creating old snapshots, creating old files.
So managing the storage in a good way means that you need to, to do it close to real time. And you also need to make sure you're not missing anything. Uh, if you check the open source, the open source doesn't necessarily always going to let you do this, even if you decide to do it on your own.
It took us a while to go from open source level to enterprise level. Uh, so on storage we can say that we actually created the most efficient storage out of those options that were evaluated. So for example, a snowflake iceberg table in that benchmark was more than two x in the amount of storage.
And this was actually smaller than a Snowflake native table in Snowflake Proprietary Storage. That's this year. So this is native Snowflake table.
This is Snowflake managed iceberg. This was a table that was, uh, managed by glue and a table that's managed by Olver, managed by Olver means that Snowflake could still read it, but Olver was managing Snowflake would call this an external table to, to Snowflake. So that's storage.
On the other side, the query performance where users get a little more sensitive. It's not just costing mon more money for the organization, we said we want the Lakehouse to perform as well as the data warehouse. So we tested it and we, I think that the, what we were most proud of in that benchmark is that we managed to get to a similar performance of Iceberg as Snowflake Native storage.
I took Snowflake, tried to query Iceberg, try to query Snowflake native storage, 7% difference. So it managed to actually deliver on the promise. Your Lakehouse is gonna be as fast as your warehouse.
And you can see that the difference from the all the other solutions was quite big. We ran this, uh, to be fair, towards the end of last year. So we are gonna run this again after we are going to release on on QTC.
The results were very good. Alright, um, switching to my last slide. Uh, these are a few examples of customers and use cases.
Mm. Uh, so there are a lot of tools in data. Sometimes it's helped to, uh, you think about what the customer actually did.
You understand what the tool is really about? I think the, the shared, uh, thread between all of those tools is you didn't have to be a specialized big data engineer in order to deploy olver. So at Proofpoint, security researchers used it to build a data lake and build software defined, uh, uh, VPN on top, uh, with many different use cases, dozens of different use cases.
You didn't have to be a specialized big data engineer at Unity. Maybe a ratio of something like five big data engineers to more than a hundred developers that took data at a scale of over 50 petabytes a year. So that's almost 50 million events per stack.
What, what does unblocked mean in that? Unblocked means that there were developers that didn't know or wanted to know Spark in order to repair data, data mesh use case in 2018 for it was called, uh, uh, data mesh, uh, in Conoco Phillips. Uh, the use case was I have an oil pipeline and I want to be able to look at every set, every me every square foot of that pipeline needs to be monitored in real time.
Other, it was, it can cost you a lot of money. And, uh, we were able to deliver that. Where the technological challenge was that I need to do a lot of merges into the data lake.
That's not, that's not a good use case for, let's call it object storage performance or to data lake tables in general. That's an, uh, something that we solved over Hive. And today we also solve other e iceberg.
I think that the co I'm repeating the comment thread. Again, you don't have to be an expert in order to do this. This is a simple user experience, but it's going to give you very high scale apps.
Flyer petabyte scale, uh, unity Petabyte Scale, do some petabyte scale, Cox Automotive Petabyte Scale repeated examples of petabyte scale use cases were one of the big reasons that Quick was interested. Yes. Great question.
Just so I understand, uh, 'cause you did talk about storage and optimized query speeds and stuff. Does this mean that, um, an Apache iceberg table has to be loaded from other sources and then we replace that old source or that source away? Or is there like a constant feed just, you know, like a constant data feed?
Oh, we got new data in from this source table, now we gotta go load this Apache table, which increases my storage footprint. Does that make sense? I'm not sure that I understood.
So I, I used to write into a regular Snowflake table. Okay. And now I want to write into an iceberg table Snowflake reads from.
So what's the difference that you're describing? How is do, But, but again, an Apache iceberg table right? Is another piece of storage that say my my SQL Server database, Right?
Yes. Have To, I still have to load this data into the Apache Iceberg table. Yes.
But I'm assuming that there's gonna be, that there is to, I'm assuming that today you have between one and three CO three four copies because you have one in your lake, and then you have one in Snowflake, and then you have one in data briefs for your, so I'm gonna combine all of this into a single copy. Okay? Okay.
Got it. Please continue. I'm actually at, uh, my last day, so before we are gonna go next, My timing's good.
Any questions? Yeah, I do have a question. Yes.
So you talk a lot about Apache, uh, iceberg, my understanding that Delta Lake and Udi are like similar to Iceberg, but different. Mm-hmm. Is that a true statement?
That's a true statement. And I would describe the years between maybe 2020 and 2024 as, uh, the, uh, as, um, format worth. So each one, each one of what you just described on the hoodie, Delta Lake and Iceberg are data lake table formats.
And Iceberg came from Netflix and Delta Lake came from Databricks and Hoodie came out of Uber. But, uh, more went more into Opensources and each one had a commercial company behind it. Databricks, uh, Delta Lake and Iceberg got the most amount of, uh, adoption Udi coming behind, and Udi was a better fit for streaming.
I think that's why it got some, uh, got some traction. The main fight was between Databricks with Delta and the rest of the market with the iceberg. And then suddenly the market and Databricks bought Tabular, which was the company that had the team behind Iceberg.
So now that Databricks are saying, instead of having a war, we are going to have both formats and try to consolidate formats and take the best of both worlds for those two format. So I think that the, in my opinion, the war is over. Iceberg won.
Everyone adapt. Iceberg in the past database did not adapt Iceberg. Now they adapt Iceberg.
Microsoft adapted Iceberg. Those were the companies that were behind Data Lake. Amazon were always around Iceberg, snowflake around Iceberg.
And Google took Iceberg as a first priority. So I would say that the, the reason we are seeing the customer running to Iceberg is because the wars are over and now it's clear what's the standard? And you can actually do something and not wait on the fence to see how the market, the market is going to lie.
I'm not saying Iceberg is better than Delta. That is better than Iceberg. You think the market chose my opinion?
Like, I'm not seeing a lot of customers are saying, Hey, maybe I won't choose Iceberg. Maybe it's not the right solution I'm gonna take, or something like that. I'm not seeing that anymore and I've been kind of asking the same questions since maybe 2020.
That's my impression. In opinion, Where do you see Click House and Starburst fitting into these kinds of companies? Click House and Starburst?
Well, Starburst is a query engine that also adapted Iceberg at the very early stages. And wherever we are going to write, they're going to be able to read and they're going to compete with the other other providers. I don't wanna, I, I have to say that for me, it doesn't matter who is going to win.
And I think the Mo more important is it's gonna be centralized storage, and then the customer doesn't have to commit to one for the next 10 years. That's the, the, the, the first thing. Click House is a little different.
Mm-hmm. Click House as a slightly different use case. We put Prietary Storage, they're adding support for Iceberg, but usually people take click outs for a slightly different use case.
So I would at the moment exclude that from that. And say, click Outs is yes, less use as a data warehouse, more as a tool that I would rep use to replace my observability solution, my Elastic, something like, like that, that I need to have, uh, um, close to real time, um, aggregated queries. We use Click House internally, by the way.
For what? For Olver logs and metrics we use in, we cannot use Olver to process olver for monitoring because we are going to, who is going to watch over us? Like, we need a, a different, a different stack for that.
So we use click Us and we like it a lot and we, it's not a use case that we would do with Snowflake and vice versa. Very cool. Thank you.
Thank you Ari. Um, hello everyone. Great to be here.
I'm Vijay. Uh, I'm product marketing lead for one of the, uh, for the data bu and, um, excited to be here. Uh, I know we are towards the fagg end of the session today.
Uh, what I wanted to, so Ari gave you a great overview of the Absorber platform, the key capabilities, some of the customer use cases, and how, you know, what type of use cases are being powered right now. Let's talk about the interesting part. For the last, I would say five to six months, we have been hard at work integrating the Solver platform click, and we have been integrating absol into our marquee data integration cloud platform called Click Talent Cloud, right?
So we are integrating UPS solver, all the goodness of UPS solver that already talked about, that'll be integrated into Click Talent Cloud. And today we are excited to introduce Qlik Open. Lakehouse.
Qlik Open Lakehouse is a new fully managed capability within Click Talent cloud that allows, that enables users to ingest, process, manage, and optimize data and for iceberg based lake houses, right? Again, as I said, we are bringing in all the goodness of UPS Solver platform into Click Talent Cloud. Again, this is not a net new product, it's a net new capability within Click Talent Cloud.
Okay? Um, not a new product, new capability. Uh, again, we are announced, we just announced this today at Click Connect.
And, um, this will be ga generally available in a couple of months in July. Okay? Okay.
So before I go deeper into click Open Lakehouse and Iceberg capabilities and what it does for customers, let me take a step back and talk a little bit about Click Talent Cloud. I don't know how many of you are already familiar with Click Talent Cloud, but I think it's good to just, to just everyone on the, everyone, everyone to be on the same page. So, click Talent Cloud today offers an end-to-end platform that delivers trusted data throughout the organization for AI as well as analytics, right?
It's already been utilized by hundreds of customers. Um, and, and it has a few key capabilities, right? Number one, um, is ingestion.
You can ingest data from a variety of sources, from a variety of sources, and then you can do either real time CDC ingestion, you can do batch ingestion, and you can land that data in a variety of target platforms, okay? Uh, it's not just ingestion though. You can also do transformations.
You can do, you can build no code, uh, visual data pipelines, end-to-end data pipelines. You can do basic transformations. You can do advanced transformations, right?
You can do, um, a push down SQL transformations and push that down to any platform of your choice, right? So a lot of customers utilize that kind of capabilities, but more importantly, you can also do data quality and governance, right? You can do, uh, you can manage data lineage.
You can even have build, you can also view trust score for each of the data so that, you know, if that data is trusted, how can I push a trusted data now into AI models? Into AI and for analytics, right? Last, but not the least, you can now create data products using, uh, click Talent cloud and push it to a variety of consumption engines, be it analytics engines, be it, uh, data science or be it gen AI applications, right?
So basically it provides an end-to-end platform that customers are utilizing to, to power their, um, AI and analytics journeys. Now, I talked about target platforms. It, it works with a variety of target platforms, including everything from AWS, uh, snowflake, Databricks, a lot of our customers use, uh, click Talent Cloud with Snowflake or Databricks or AWS, Microsoft Azure, you know, and Google Cloud.
Now, if you think about it, iceberg becomes a net new target platform, right? Hey, do I wanna push my data into Snowflake, or do I wanna push my data into, uh, into Iceberg now? So as you see here, we are now introducing Click Talent Cloud, or click Open Lakehouse as an integral part of Click Talent Cloud, um, that enables ingestion into iceberg optimization of Iceberg so that you can drive all of this data for analytics and data science use cases.
Okay? And, and I'll double click a little bit more on this, but I want to give you an overview of Click Talent Cloud before we dive deeper, deeper into the, into the iceberg capabilities. Okay?
Now that we talked about click Talent cloud, let's talk a little bit more about, let's double click on click open Lakehouse, right? The iceberg capability that we just talked about, talked about. So if you're talking about, um, click open lakehouse, there are three key things or three key capabilities that I, we would, we should highlight.
Number one is the high throughput ingestion. And Orik kind of mentioned this alluded to as part of the Solver platform. What you can now do is you can now ingest data, real-time data, or batch data from hundreds of sources directly into iceberg tables, into query ready iceberg tables on Amazon S3, right?
So you can do ingest data, um, uh, in, in a batch fashion or real time fashion, and use that, use that for analytics or for, um, for machine learning purposes. That's number one. Number two, adaptive Iceberg optimizer.
Ari alluded to this again, as part of this is in my mind, one of the key capabilities and differentiators that absol or rings to the table. It's the ability to, um, do always on optimizations, right? You can, it, it monitors iceberg tables, and it does, and it determines the optimizations, cleanups, compactions required to minimize the storage footprint and improve query performance.
So I think that's the, that's the critical piece here, is the adaptive Iceberg Optimizer, that that's part of, part of Click Talent cloud. Now, last but not the least, another critical capability is what I call data warehouse mirroring. Now, what this does is you can now mirror data from your iceberg tables directly into Snowflake, right?
You can mirror this data directly into Snowflake, um, for downstream transformations. By the way, you can do that without duplicating the data or creating another copy of the data, right? So, which means you can now run your ingestion and bronze layer on iceberg tables, but then mirror the data for silver, gold, and any of the other more curated data sets, right?
So those are the three big attributes. But one of the other key things as part of the launch that we are announcing is the integration with catalogs, right? Apache Iceberg catalogs, we will be integrating with three key catalogs, including AWS Glue, Apache, Polaris, and Snowflake Open Catalog.
Now, what does that mean? The moment you are, you are, you are now integrated with this catalogs. What that means is now you are opening up the data to a whole set of query engines, right?
You can use Snowflake, you can use Amazon Athena, you can use Spark, Apache Spark, you can use Trino, Presto, Dremio, any engine. So this is the promise and potential that Iceberg brings to table, right? You store data once, but then, then you can utilize the data across a variety of engines and processing engines and query engines, right?
That's the promise of, uh, Apache Iceberg, and that's what Qlik Open Lakehouse brings to the table. Can you specify, when you say high throughput, how high throughput? Yeah, so, um, Ari talked about it a little bit, but some of their customers are bringing in data as much as like five to 50 million even per second.
So, so again, pretty high. That is pretty high, right? So, because they're talking about use cases like cybersecurity, uh, iot, right?
Where you are constantly looking at streaming data, um, coming in, in high volumes, right? Again, now the, so streaming ingestion is part of it. You can do real time ingestion from.
So that's actually a good question because it's a good segue into my next, um, storyline around ingestion, right? So what you can do with, um, with, with, um, click open Lakehouse is now you can ingest data from any of these sources. You can do real time CDC ingestion, you can do batch ingestion from cloud sources, from, um, SaaS applications, from file sources, from, uh, any of the operational databases, right?
Uh, Oracles or any of the operation, MySQL, any of this operational databases, SAP mainframes. So imagine you now can ingest our data from SAP and have it available in S3 tables and then open it up to a variety of query engines and processing engines. It just opens up a world of possibility and opportunities, uh, for our customers, right?
So, um, so now you, you can, you can bring in data from any of these sources, again, as I mentioned, um, CDCR batch. But what, what, um, click Open Lakehouse does is it handles the hard parts very, very easily. For example, it does, it automatically maps source to target data types.
It resolves type conflicts automatically resolves type conflicts. Um, it handles schema evolution and data drift, right? It, you can easily add or delete rows, um, uh, using, using this.
But, but the, the, the critical aspect here though is you can do all of this without the need for a data warehouse. You don't have to burn data warehouse compute for ingestion or for the bronze layer, right? We make it really, really simple, easy and efficient for customers to bring in data from any of these hundreds of sources into Apache iceberg in a very seem seamless and easy fashion.
Okay? So that's kind of the first attribute that we talked about is around ingestion. Um, um, and Antoine will talk a little bit more about the roadmap.
Uh, what do we have coming up today versus what do we have coming up in the future? So it's a lot more interesting stuff coming up in, in terms of ingestion in the near future. Okay?
So we talked about ingestion. One of the other attributes I talked about earlier, and I think already touched upon this in my mind, is the secret sauce from UPS solvers really is this adaptive iceberg optimizer. Right?
Now, think about it. If you are building a lake house, iceberg waste lake house, you are worried about two things. Number one, you wanna make sure your storage footprint and costs are minimal, but at the same time, but not at the cost of query performance, you wanna make sure you have very comparable, very good query performance as well, right?
Low latency query performance as well, not at the cost of each other. You wanna maintain low storage footprint at the same time, maintain a very, very competitive, very low, uh, query performance as well, a very low, uh, query latency as well. Um, but, but think about it, what you don't wanna do is have a team of 15 engineers who all they do is sit and tweak and tune every single iceberg table day in, day out.
That's not gonna be feasible, that's not gonna be, not gonna be scalable. That's the reason why, um, why absorb build this click adaptive iceberg Optimizer, right? And what it does is it is intelligent, it is always on, you don't have to schedule these optimizations.
You don't have to do these optimization session manually. It continuously monitors tables and thus optimizations, compactions cleanups dynamically to make sure you have the lowest storage footprint. And, uh, a, a very competitive, uh, uh, very compelling query performance as well.
So, so for example, like the cost-based compactions, it dynamically compacts files based on a number of variables, including, um, number of files, um, um, file sizes, uh, frequency, frequency of updates of these files, right? So a number of parameters are taken into consideration to do these compactions in a very, very automated fashion dynamic partitioning, you don't have to worry about partitioning the data or figuring out testing and planning out table layout implementations. It does it for you, and it determines the ideal table layout that might work for your data type.
Yep. So, On another slide, I thought the, I, I saw that these are all workloads, right? These are built-in workloads.
Do, are they customizable or are these just, uh, do you pick the ones you want to run? Are these just, it does say always on. So are they always on whether you want them to be on or not?
Maybe there's some things you wanna ama you wanna, like, you do wanna manually tune. Yeah. So, Ari, is there a capability for them to actually pick and choose some of them?
And then actually manually, Uh, what I said is that every observer or click job is attached to a cluster. So once I shut down that cluster or stop the job, it'll stop running, which I can do via UI or programmatically. So the default is always on, which is what, usually what we do with data integration.
Uh, you can decide not to. I'm not seeing a lot of use cases for that, by the way. Like, usually people just want their, they want their data all the time.
They want it fresh. Awesome. All right, I'll keep that.
Um, and, and then the last last thing is about intelligent cleanups. I think, um, kind of already kind alluded to this. It's about, um, snapshot expiration, right?
Deletion of all these or orphan files in a very safe and consistent manner. Again, your objective is to minimize the storage footprint, improve query performance, right? So what does all this mean for customers?
What what does this result in is, you know, one, obviously you get to ingest fresh data, right? From all of these diverse sources. You have access to fresh data within seconds, and you can analyze that in a variety of ways.
Obviously that's important. Number two is, you can, because of all these optimizations, cleanups, and compactions, you are now able to reduce the storage footprint and cost by up to 50%. And you saw some of the benchmarks that were shared earlier around storage performance and query performance, right?
Storage performance improvement, but not at the cost of query performance. That's what we are aiming for, and that's what, that's what customers want, right? So you are, you are talking about query performance comparable to an experience that you can expect from our data warehouse, right?
That's what customers can, customers can really, uh, achieve, uh, from click open Lake House. Um, with that, with that, I wanna, again, zoom back, zoom out a little bit, right? So we were digging into the open lake house, but keep in mind this is a part of a bigger platform.
The end-to-end click talent cloud platform, which allows you to ingest or any types of data you can now transform, you know, data transformations. You can do data governance and quality, right? And open Lakehouse fits right into, uh, into click Talent cloud, where you can store ingest data directly into iceberg tables.
You can do compactions, um, you can do, you can do cleanups, and then you can query using any number of engines as we talked about, right? And serve that data for analytics, for building data products, or for building Gen AI applications. So that's what an end-to-end QTC or click Talent Cloud offers customers.
So it becomes an additional compelling reason for customers to adopt QTC. Okay? Now, um, the beauty of all of this is, and again, we are just launching this now, and, um, all, so click Talent Cloud has four, um, pricing additions in terms of pricing and packaging.
Everything that we talked about in terms of, um, click open lakehouse is, will be available from standard addition onwards, right? So, so all of the iceberg capabilities that we talked about, including ingestion, including optimization, uh, including, you know, opening up all the query engines, all of these data warehouse monitoring, all of that is gonna be available starting from the tradition onwards upward. So we are making it really, really easy for customers to get started, get started building their open lake houses, get started ingesting data into iceberg and power, some of these very compelling use cases that you, that you saw.
Okay? Um, so I think I talked a lot about the capabilities and, and how this kind of all fits in. Uh, but, uh, you know, again, a lot of what we talked about now is already launching today.
Um, but we have more interesting stuff coming up in the roadmap, right? So what do we have now? What do we have coming up later?
I wanna call on Antoine, uh, who's gonna share a little bit more of the roadmap, but also show you a demo. So, so, so more on that pretty soon. So, moving on to this map slide and talking about, uh, timelines here.
So in May at Connect, we're just launching the click open lakehouse, uh, and we are target targeting a July for, sorry, a GA durability for July. Uh, what will this GA include? So we'll support ingestion from all the click talent cloud sources.
So that's all the sources that vj vj, uh, listed before. So that includes SAP databases, a lot of SaaS, uh, applications, as well as your mainframe system. So all of these capabilities, the CDC and Batch are existing in Clicktel Link Cloud, and we plug that into the UPS solver technology from the ga.
Then obviously we include, uh, the UPS solver technology that's, uh, with the Adaptive Iceberg Optimizer and, uh, the ability to store data as iceberg tables on S3 and publishing to the catalog that were named before, the Blue Catalog, Polaris, and the Snowflake Open Catalog, that's obviously a no data warehouse, uh, tion system and storing system, we have as well. With that, a mirror into Snowflake, and I'll dig more into that. We have the ability to make this data available to Snowflake consumers for BI use cases.
Typically, I, I'll dig more into that, and this will be AWS only at launch. We'll add more cloud providers later on. So, I have a question.
How, how seamless will this be for customers who are on, like, for example, standard tier of QTC? How much of it is going to be like a faceless, it's all happening under the hood, and how much of it is going to be going through a wizard? And how much of it is going to be, be I have to change my setup?
What does GA actually mean from that perspective? So in term of the user experience, that will not change for, uh, users, that's using QTC. Today we have, we have the storage capability and the project capability that's tied to a target platform, snowflake, Databricks, and so on.
We are basically adding another platform, which is the clear open lake. So we'll start a new project that will work exactly as a Snowflake project, but on Iceberg. So in terms of map, we will be adding, uh, streaming capabilities.
So that's the olr technology streaming into Iceberg that, uh, allows customer to ingest data at fast paced and, uh, high data volume. We'll be coming soon by the end of the year. Uh, as well, we, we want to support our customers who are using and running click Applicate, OnPrem and, and click, sorry, and Talent Cloud.
So we'll work on ridges, uh, so that those customers can leverage the benefits of the clickable data balance integration. So, as I said, we have mirroring to Snowflake in the first version. We want to support mirroring in data risk Redshift and others.
After that, more cloud wire supports Azure first, so AWS ga Azure, then, and then GCP streaming transformations and data quality expectations. Those are the, uh, capabilities that or mentioned that allows us to, to do, uh, transformation on streaming data as data flows in and creating other ice variables, other iceberg objects that are the output of the transformations. So that's basically, uh, the three key capabilities of olver that will be integrated in Click Talent Cloud.
So now looking at, uh, a typical Click Talent cloud project, that's what our customers do today. So this is a screenshot of a pipeline in Click Talent Cloud. On the left you have your data sources, and we have our CDC or batch ingestions that is loading data to what we call a landing area.
That's the raw data. This is here to support the CDC process. We land the data row as we replicate it, and then we take it to this bronze layer.
This is our storage task that takes care of doing the merges, takes care of making the data available for long term. And we, we typically create a type one and type two view. That's SCD.
So type one is the current version of your data. Type two is the ized version of your data to go back in time, okay? And then you have, you can have more layers, typically silver angle layers, you can, you have, we have specialized tasks that allows you to transform data and make that fit for your, uh, downstream use cases.
This is the type of project that you can create today with Snowflake, Databricks, Redshift to whatever. What we are launching in July is the ability to have your bronze layer, your storage layer with type one, type two on iceberg. Okay?
So in term of user experience as the same kind of project, it's just another target platform. What we want to do as well is we want to look to store data on Iceberg, make that available through, uh, iceberg catalogs and as well to their data warehouse platform. They all use one or many data warehouse today.
We want to keep that stream, uh, seamless for them. And for example, uh, ability to mirror the iceberg data to, uh, the Snowflake database. Okay, so I will show, I will show a demo that is a recorded demo.
Uh, I've done that a couple days ago, and I will comment on it live. So this demo will feature, uh, data es from a SQL Server database into Iceberg Data. This iceberg data will be published to the LLS Glue catalog and consumed by the Ater Query engine.
Then we'll have a second part where we'll mirror the data to Snowflake and have, uh, transformation capabilities on Snowflake, uh, showing you that you can consume and read iceberg data from Snowflake. Okay, let's go. So opening one project here.
So this is an example pipeline from the source on the left and going through the different tasks here to the right. This is a project that's created on the clickable lakehouse target. I open the settings, as you can see, platform type here.
This is where you can have Databricks Snowflake. We have the open Lakehouse, we have Configuration to Blue Catalog and two S3. So we are replicating data here in CDC to S3, thanks to this, uh, what we call Lake lending capability.
That's storing the raw data and inserting all the changes as they are made in the source system. We insert that into S3 in this, uh, particular, uh, director here. So that's S3.
This is where we store data or whole data with our lending capability. What's interesting then is the storage task Storage task exists today. This is a storage task that runs on Iceberg.
So we are applicating taking care of these data sets here, as you can see on the left. And as part of the settings, we have as well the directory on the three when we store where we store data. So this is the other directory that I have here.
And as you can see, these are all the tables that I'm replicating from the source. These tables are registered on AWS glue in the, in the glue catalog, under our database name, and are available to any is BR compatible query engine to query. So here I'm showing you an, an example of a query using AWS aina and creating the tables here.
As I say, we'll support, uh, type one and type two. These are views, uh, for the type one view we are currently working on the, on the type two that will be available for GA in the summer, creating type one version of the data through ATR or any, uh, iceberg compatible query engine. You can also see the physical objects, meaning the iceberg cables through our design ui.
That's easy for, uh, designing and, and troubleshooting. You can view the data as well, iceberg data directly from the, from the tool. Then I will show the mirroring to Snowflake.
So we've added a new, oh, sorry, the Lakehouse cluster. This runs with the Lakehouse cluster. We have a compute technology that comes from Olver that we deploy next to a three on the customer infrastructure.
And that's a number of nodes and the customer scales up and down based on demand and based on the data volumes that need to be Es you have, uh, scaling strategies as well that you can define, uh, if you want to control cost or as opposed to that if you want a lower latency. So jumping into that, uh, mirror, uh, task here, that's a new task that we've added to the product that's, that allows to mirror ice to another platform here, snowflake. So we have a configuration to, to Snowflake here and a set, uh, of settings to, so that Snowflake can access the escalator, and Snowflake will connect to the Glue catalog as well.
So what's interesting is that what we create on Snowflake, so we have, we create this more task, creates a specific schema, and I'll show you these external, external tables as they call that. So that's external table relative to Snowflake. So this is a nice back table here that's rated on Snowflake, using an external volume and a catalog that we manage with click Open Lake House.
And we issue refresh command to Snowflake to make sure, make sure Snowflake always has, has the latest data and metadata. Then since it's more to Snowflake, I can create another project using Click Talent cloud, which will be a snowflake project in which I will consume this sizeable. So this is a snowflake project.
As you can see here. This exists today. And as part of that, you can create a transform task that will take data and transform them.
This transform task here is using iceberg data from our iceberg, uh, the former Iceberg Openhouse project as a a source, as an example. It takes two data sets and transform that using, using the transformation flow capability. This is very similar to the data flow capabilities that were was shown before.
This is for, uh, SQL Pushdown. This is a version for SQL Pushdown, and we generate, uh, SQL commands based on the visual transformation here. So you have two data sets, a join and aggregate.
We generate an output table that that's translating to a, to a, to a SQL segment automatically for the user. And we push that down to the target platform in which, in this case, uh, snowflake. So this is a table that we've created, actually, you can configure if whether it's a view or a table.
In this case it's a view here. So e every time you query this view, that will actually query the actual iceberg tables showing that on Snowflake to show you that the object is actually created on Snowflake. This is, this is the view, and if you create, we can, uh, provide the, the output.
So that's it for the, for the demo. So new capability as part of AL Cloud. And as you've seen, not a lot of change in term, in term of user experience.
That is all for, for us. So, quick recap here. Uh, high volume data digestion, hundreds of sources.
All the sources that we support today with AL Cloud will be, uh, ingestible into iceberg, will add streaming, uh, streaming sources as well, very, very soon. Huge cost savings. Uh, you don't need a data warehouse to ingest data and store data for your bronze layer.
And we have the Iceberg Optimizer as well. All of that integrated, uh, with the Click Talent Cloud platform.