Evan Kaplan on InfluxDB: Time Series Data, Open Source, and Real-Time Applications | AWS re:Invent 2025
InfluxData CEO Evan Kaplan discusses InfluxDB as a leading time series database, emphasizing open-source community support, deterministic data collection, and its integration with AWS. He highlights challenges in the competitive database landscape, the growing role of data scientists and engineers, and encourages developers to engage with time series data.
Transcript
Hey everyone. Welcome back. We're here at AWS reinventing our suite up at the wind, continuing our coverage of this year's event.
This gentleman, am I right? I actually, I think the last time I saw him was at Reinvent. I don't even think it was last year.
Evan, I think been two years, I think was the year before. Yeah. What a great, but he's a, a wealth of information and, you know, there's six degrees of separation in tech.
He's probably his three degrees of separation with this guy. He, everyone I know knows Evan Kaplan. Oh, geez.
You run in some bad circles. Yeah, I guess I get that's maybe what it means. All right.
Evan, of course, is the longtime CEO of Influx database. If you don't know influx database, shame on you, but we'll, we'll tell you about him in a second. Evan, it's good to see you, first of all, Matt.
Nice. Alan, thanks for having me. A pleasure.
Um, you know what, let's start with Influx database. I, look, I'm going to guess 80, 85% of our audience may even be using influx at some point, if not now, some High percentage. Yeah.
Well it's, you know, our people are your people. They're, they're, they're type people, right? So, but for those who maybe aren't, they don't, they've never heard the word Time series database even Explain to them, if you can, what, what we're talking about.
Sure, Sure. So, first of all, InfluxDB is an open source time series database really is the first open source time series database. And, and by all external measures a leader by a big margin in the space.
Mm-hmm. 4 plus million people who run it daily million sites run it daily from the largest corporation down to, you know, down to people using it at home and that sort of stuff. Mm-hmm.
And so our, our primary, let's call it invention or innovation, was by my partner Paul Dix, who in 2013 made the first commit. He had worked on Wall Street and had built a number of time series databases on top of other platforms, HBase, sql, things like that. And, and it turns out that there's enough difference, enough handling enough optimizations that are available, um, when you're working with this kind of data that it made sense for it to be one of those categorical databases like search, like document like graph.
Mm-hmm. And so he's generally considered the person who sort of invented this space, a modern version. The first version of the product was written in Go, it had collectors now a project called Telegraph, which is actually more, more popular than the database.
Really? Yeah. They're about 400 different collectors for virtually every system out there, physical and virtual, you can imagine.
And so it's a super vibrant community. We've stayed with a promiscuous MIT Apache license model, so anybody could use That. Use it and do.
Yeah. Yeah. And if it's worth, I mean, just as a refresher for the audience, why, what, why, what time series is, why is it, What do we mean, Right?
Why is it? And so really it's quite simple idea, which is anytime your primary tag, anytime your primary index is gonna be based on time, you can be incredibly efficient at handling large volumes at super fast. Um, actual real time kind of, um, kind of, um, querying and ingesting and things like that.
And so if, you know, you're gonna have that kind of application, IOT physical AI network telemetry, where you're gonna need that kind of speed, this kind of huge ingest, you start by building on the crime series specific platform. So you get away from these very discreet indexes that take time. You get away from these latencies associated with queries that have to be, you know, aggregated in different fashions.
You also have to do a bunch of things. You wouldn't, you have to downsample you start with high resolution, maybe even millisecond resolution, and maybe you wanna downs sample to a minute. You have to evict a lot of data.
So just a bunch of stuff that you would do that a normal database doesn't do. So the category has really emerged strongly. Yeah.
So we're not the only player. We have good competitors. Azure offers something now Amazon offers us, like, so it's a market that's taken off.
Alright, that's too long. But No, no, it's not too long. You know what, for people who don't know what a time series database is that I think it's necessary.
Look, for me personally, when I hear time series, I think influx. I think because you had the open source model and, and the commercial offering, I, you know, there, I know there are other time series ones, but other time series databases. But, you know, the database market isn't what the database market was when I was much younger.
Right. And, and we do have, we have graph databases and vectors and this, but when, when Paul Dixon did this in 20 20 13, He made his first Command. He wasn't thinking about, well, how's this gonna work with generator of ai?
Right. He did have, his background was in machine learning, really. He wasn't completely blind to it, but no, he wasn't thinking that.
And correct me if I'm wrong though, this is really the, the whole AI thing. Well, AI's helped a lot of the data, data collection, data storage, you know, access data's a big part of it, but it, it's been a particular boon for, for time series as well though, right. There's some real use cases.
Yeah. A certain, a certain, a certain kind. So, um, if you just sort of, you know, you break the world into, and it's increasingly become apparent, you have generative ai, which is largely the scraping of the digital world.
Yep. So anything, you know, and now we've scraped pretty much everything. It's available and every day we scrape it again.
And so there is a real shortage of digital, digital data, images, all that sort of stuff, because most of it's been scraped in some sort of way, scraped clean, and it's, and it's, and it gets, and it gets indexed every, you know, every day. That is not, and so the issue on, on the generative side is you have to create synthetic data in order to build these models. We're To make more sophisticated Right.
And that sort of stuff. But the opposite is true on the physical side, right? If you think about the physical world, the amount of data available to index is, is, is infinite.
Right? It's just a question of what resolution you want to capture it, how often you want to capture it, and that sort of stuff. And so what what we're most excited about is, you know, as 'cause of our IOT orientation is the physical AI world.
Because, you know, the, the notion in physical ai, the idea is I want to get a near perfect picture of the physical world to build my models around and to act on that world, right. To build my intelligence around. And so the only way you get that is by willing to take sample measurements that are really, really close Time intervals.
Yeah. Yeah. So our default, our default timestamp is nanosecond.
Really? Yeah. And so we have some customers in quantum world who use that, but most, most, you know, you're not nanoseconds.
No, but you know what, there was, I forget his name right now, I apologize, but he was like the chief AI scientist or one of the chief guys that meta who just left I know if John Koon. Yes. Yeah.
And he said, because we're probably coming to it, I don't wanna say a dead end, but a, a diminishing returns because of the problem you, you've outlined here. Right. Which is we can scrape, there's not much left to scrape.
And that the LLM model only works. It, it, it can't imagine a sphere in open space and how it's gonna fall with gravity. Well, let's say right?
Where, so he calls them world models, right? Not LLMs, but, and they try to model digital models, the physical world. Yeah.
And, and that's the kind of database, you know, data collection that's gonna be huge. I mean, I, I mean that's a really, that's what we're most excited about. And if you think about, you know, LLMs are probabilistic models, right?
Yep. Right. And so, and they're brilliantly probabilistic.
Right. They're actually really amazing. Um, but you, you know, when you, when you're operating in the physical world, you think of a robot, a self-driving car, a satellite in space, you need deterministic.
You cannot live with probabilistic. The chance of your robot, you know, accidentally vacuuming your cat away is just not something you will, you wanna live with. It is not something you wanna live With's not, that's not risk I can manage.
And so, and you said this, so the notion of deterministic versus is probabilistic is important. And so collecting this data, building your models on really rich physical data, and then being able to act in a control system in a meaningful way, that's the future. Absolutely.
So, hey, we could talk more about that if you want, but I, I gotta bring us back to today if Yeah, sure. Okay. Um, so there is the open source, then there's the commercial version.
You guys offer like a SaaS version of it as Well? Yep. We have a couple of cloud versions.
Yep. And we have you, you can run it OnPrem or run it in your own instance, And you could run it in the cloud. I know you've had a longstanding relationship with Amazon with AWS but you guys recently did an announcement, maybe it was a month or two ago already.
Yeah, Yeah, yeah. We have, um, credit to AWS we have a, um, we have a very unique relationship with, so we were always a marketplace kind of. Yes, you could, you could purchase influx.
We run our own cloud services, our serverless platform on Amazon. Mm-hmm. And so we've always lived in that ecosystem.
We've always worked in that ecosystem. Um, but what's unique is about two years ago, they reached out to us and, um, they were getting a lot of requests because we have all these open source influx tvb to run influx on their platform. And so, you know, immediately I was like, you know, that doesn't sound great.
We're an open source player. They could just take the code. Well, it would, And it's not like it hadn't happened Before.
I was just gonna say there is a reputation there, but go ahead. And so, so I was, I was a little bit like, okay, let's, let's, let's build some trust here. And, um, and the relationship we have is the first time they've, they've ever done this is they licensed our open source code.
They run our open source code, and then we build the value added enterprise features on top of it. And we monetize that, not in marketplace, but in console. So when a customer configures, I love it.
A certain configuration, we get revenue AWS Correct. So you'd ask why would AWS do that? Um, you'd have to get Brad, which I'm sure you could, who runs, runs those businesses mm-hmm.
On, but, but a couple of things. One is they saw the emergence of time series data, and they do have a platform, but it's, it's of a, it's a pretty, it's a narrower use case and influx. And they were getting a lot.
And so just recently they didn't, they stopped taking new customers for that platform, Their end of life. So their primary time series is gonna be time stream for InfluxDB. And so in their world, that's a one P product.
It's a product, their product, it's not Influxes product. It's not their Product. It's based, they we work with And we are close with them, and we, we work hand in glove on those deals.
And we're more than happy when customers wanna, wanna purchase that, go Through through AWS That works for us. It works For that. It's a great channel.
All That sort of stuff. Let me stop you here for A second. Go for my people watching this at home.
I, if you are not in this side of the business, if you're a, a consumer of this Yeah, sounds great. They got a good relationship for people on this side of the road though. I, I bet you you could count those kinds of relationships with AWS on one hand, you Can't even count on one.
I think we're the only one now. I think they will, I think they will do more because it's worked well with us, right. Because they don't wanna work this stuff.
And unique Yeah. You had a unique situation where, hey, they don't want to use you. Good luck.
Go use what you want. You know what I mean? And, and there was no choice.
Right? Uh, influx is the bomb as they, you know, the shizzle. Oh, I'd like to think we have that kind of market power.
I don't think we do. Oh, there are, but I'll tell you, as someone who sits here, I don't have a in this race. Okay.
I appreciate that. It's kind of like secretary of, we certainly don't Operate, we certainly don't operate without Mindset. Don't let it go to your head.
Stay humble. Stay humble. But when I see, I mean, it's few and far between that we hear of another time series data.
The hyperscalers have their in-house one, but you know, that's the old 80 20 rule. Yeah. Say 80 20 rule.
That's exactly Right. It's 80% of the function, 20% of the price. They're good if you wanna hit the easy button.
And so if you decide you really wanna Build something, but if you're serious about time series data. Yeah. I, I get it.
I guess though, it begs the question, what about the other hyperscalers? What about some of the other providers? Well, it's, um, we still run, we run our serverless platform on Azure and Google.
Okay. So those are available, um, it's likely what the announcement of last week was, not the, the relationship, it was the, the introduction of our version three into the Amazon. Oh, okay.
So that happened a month and a half ago. And that's we're most excited about. That's our canonical version.
Now. That's where over time we'd love everybody to land on this version three, which is a almost a complete rewrite of is it? Of all, yeah.
Of everything we've learned over the past 10 years. Let's, let's hear about some of the, Yeah. So version three, killer Features here.
Yeah. So version three, um, we just ran into a number of issues as, as customers scaled on our older versions. One is cardinality, and I know you're familiar with this, but maybe your listeners aren't.
When you get, you know, when you get the exponential, you know, the number of tags and the number of descriptors exploding, it can really choke off a time series database. It doesn't, it's not much of a problem in observability world. But in iot world, it's a pretty big problem because you can have tremendous variations.
Sure. And so cardinality started to choke off on older databases on that model. Um, compute and storage were linked together.
So that was expensive as you scaled, right? Most, and most databases are that way. But, but, um, unlinking them, we didn't support Native SQL before, really.
We sorted influx ql, which was SQL-like, so you could use it, but like all the business applications, the other stuff that were Needed, quel Q sql, you had to, and The tools that were built SQL weren't, weren't, weren't optimal. Um, and so we, we ran into those problems. And so we, we started writing this about four years ago.
We wrote it in the open source, and we wrote it in Rust. Really, The original database was in Go, we, one of the early offerings With Docker and go, but we in Rust. And, and Paul and the team had decided that Rust was gonna be a much stronger platform over the future.
More secure For memory, safe, more secure, just a better platform for running database. So that was a rewriting rust. And then we built it on object storage as the date of storage, which, so you could have long-term retention without worrying About.
That's great. And then we obviously ordered sql, and then we did some really other interesting things. So we scaled it differently.
We scaled it horizontally. We, we always scaled horizontally. We just scaled it differently.
So it was much easier to duplicate. Mm-hmm. Um, and then probably most importantly, we built something, a processing engine directly into the database Right in, Right into the database.
So there are probably six or seven triggers in the database that the processing engine can call. And anybody can write a Python application script or anything to use those triggers. No kidding.
So if you want ETL stuff on the fly, if you want to transform stuff, if you want to down sample stuff, if you wanna do anomaly detection, although all of a sudden developers can write their own stuff and actually exercise it in real time at the same, at the same latency, the same performance that any database query will. So this begs the question, can I set up like a Python script repo for influx for these things? We, We have a, we have one of that in GitHub, right?
Okay. We don't have an, well, what you'd see as a traditional app store, but in GitHub, we already have a repo and people are building stuff. I mean, it's pretty exciting.
So the idea is that that helps us move up stack. And so you wanna do forecasting, you wanna do anomaly detection, you build it Yourself. You want to do, have To buy third party Pro, right?
In today's world, it's all about scale. So you took a 2013 database and brought it into 20, 25. Painfully, just wanna say painfully, Well, none.
You know, if, if it, if no pain, no gain. Well, It turns out, you know, one of the things, you know, I came from, um, from the networking and security world, not necessarily from the database world when started here, one of the things you realize is like, you cannot accelerate the time it takes to mature a database. No.
Like, you can't just say like, I'm gonna have a database built in a year and think that it's gonna be able to scale. You're not gonna have gigantic. Like it has to run in the real world.
So in some ways, now that you look at ai, we're, we're kind of the boring part of that infrastructure, which is fine, but we're the least disruptable part of that Infrastructure. Right. And, and, and it's nuts and bolts kind of stuff.
Unique thing here though is you have the open source. How big a help was the community in this kind re relaunch? There are places where the community was Actually, excuse, I don't wanna call it a relaunch, Evan.
That's a big Word. Yeah, no thanks. Yeah.
Yeah. Just a, in This, this the Version next gen, the next generation version. Um, the community is super helpful in certain areas.
It's really hard for the community. Really helpful in the database. But, but here's the distinction.
We built this based on a bunch of Apache standards, some of which we're the main committers. So Apache Data Fusion is and Apache Arrow, the, the, the memory format. So we now commit to those and we use that in it.
So we took advantage of the community. And then data fusion, that's our, you know, that's a project where we're the PMC. It's a super popular, it competes with, you know, with with the query planners and engines that are in Databricks and, and in Apple's, proton, things like that.
Mm-hmm. And so that's a big part. So we leverage those.
And then we're really also leveraged on the connectors. Our community's constantly writing these telegraph plugins. And that's the bread and butter here, man.
You Gotta be connected. Yeah. Yeah.
And you gotta be connected without ETL. Right. And you can do that in time series.
What about, uh, MCP servers, stuff like that. Yeah. So we check, yeah, I mean, sort of check four months ago.
We are now using it in version three, so that people can do natural language queries against, um, against, directly against the database without knowing SQL or that sort of stuff. Some customers are using it. I think it's really important.
It's just, it's not driving our business today. Yeah. But it's really important.
It is. Um, so I, I gotta imagine with Amazon's, with this deal with Amazon and version three out here, you, it must be attracting a decent, uh, amount of buzz at the show. It's, you know, this show is Crazy.
I know, but it is 60,000 people basically. I mean, this show is, yeah, I don't know what was a decent buzz at the show? Their keynotes are the word casual buzz at the Yeah.
At the show. But no, it's in general, it's, we're just in a really good time in the business. Having the database out hit, having, you know, the level of maturity that we finally need, Amazon coming in, it just feels like a bunch of forces are working.
The narrative now in the community is really much stronger around physical ai and people are paying a lot of attention to that kind of work. And so, you know, but, but this event, it's, you know, I think, I dunno if you read the keynote, but, you know, I talked to people who were, it's like everything was about agent, agent Ai, it, that it was all agent, agent ai, all Database stuff and all the infrastructure, all the cloud Stuff. They didn't talk about data at all.
Forget database. They didn't even talk about data used to come to, uh, to this show. They talk about that three and, and Lambda and Serverless.
Yeah. I it's still there if you dig deep enough. Right.
It's not in the main keynote. Not in the keynote. And I get why?
No, I mean, I understand. I I, so we've spoken about this on our shows. Look, I personally think they needed to show they can go ahead to toe to toe with Google at Microsoft on ai.
'cause Google has a good AI stack. Yeah. A hundred percent.
A hundred percent Theran processors, all these things. And For the last 10 years, the narrative is about cloud computing. And so, and now, you know, cloud computing's default now, Right?
They don't even, it's Reach cloud here. Yeah. It's all, you're right, it is all ai.
Um, I I, I had a question in my mind and I got stuck on this thing with you right here, but I'm sure it'll come to back to me in a second. But let me ask you, actually, I do remember it. We have, you mentioned graph databases before.
That's huge. Bigger than it's ever been, right? Yep.
Yep. And, and, and the other thing, what's driving a lot of this data is this whole observability space. I mean, you know, it's hard 'cause AI sucks the oxygen out of every conversation we have.
But you look at like observability, you look at, you know, the growth in graph databases. I've spoken to a lot of end user organizations where they're using a time series database. They're also using Graph, they're also using sql.
Right? It's, it's not a one database town no more. No.
I mean, no. In fact, you know, my take on this, and you, you and I are of a similar vintage is, you know, when we were growing up in the industry, at least in the early years, there were only two databases that matter. Yep.
It was, you know, Oracle and IBM PB two. Yep. And so you just choose which one, you know, and then we saw the, the emergence of MySQL and the open source, and that, that stuff probably happened, started, you know, 18 years ago or so.
And then you saw the, the taxonomy of dropping into these specialty databases, right. Documents, seql, sql, and all of the no SQL graph and all that. And now those are pretty well-defined categories with, with large competitors.
And so, you know, it really, and the ability to assemble the appropriate data models and do it are important. But I think the important thing is is, or the Oracle and IBM of this world that I foresee, you know, are the Databricks and the snowflakes, and now the fabrics and the big queries and The, they, they are, Those are the databases, right? Yeah.
We are all, we are all live in that constellation. Our data, if done right, the models are built somewhere in Databricks or Snowflake, right? Yeah.
There, all the structured data, the unstructured data, it's all combined. The intelligence is built there. And we view ourselves, and this is distinct, and most people don't dive in, not as an analytical database.
Even though our analytics are great, we view ourselves as an operational database, really. Yeah. We want people to build control systems.
We wanna build, you know, dashboard, list monitoring systems. We wanna build automation. We want people to use our stuff And integrate to be active.
Right. Right. We don't, I mean, well, because It's analytics Are there.
Yeah, no, I get it. I mean, you have to be, it's table stakes. Yeah.
But that's an, I never thought of it that way. That's actually a good, So you think about our customers, you know, whether it's Tesla trading energy from the ba, the power walls, whether it's, um, Utah sat, um, or Kuper doing, you know, the early satellite positioning and the change of positioning. Mm-hmm.
Like these are all things that are happening in somewhat real time. These are things that Lakehouse can't do. Yeah.
Right. They're latency. The orientation, the cloud.
Only half of our customers are three tier architectures where they have stuff on-prem and in the cloud. Like you need, like, it's totally different. It's a very different world, Different thing.
You know, it's this whole, the, the, the data, the age of the data scientists, right. And, and this AI thing has giving them rocket fuel. Yeah.
Yeah. Right. com.
So we're very big in platform, right? Yeah. The biggest element joining that community data scientist, it don't make sense to me at some level, but data scientists are very interested in making sure they have this platform.
So where, So usually you say it that way. So what, what what we see going on, and it's early, is that there's emerging category called the, you know, that that feeds the data scientists, that feeds these data engineers. Yes.
Right? Literally, they're not sre, they're not dev, No, no s they're Nothing. Right?
Right. These are people who build these pipelines. So you have these distinct database, you have these flows.
They own the engineering. They're like, if you think of it, I mean, we used to talk about plumbing. They own the platform.
They're the real plum, right? They own the Platform. How does the data move across these applications?
How are they made available to the model? How do they operate in real time? That if I could train, you know, I have two college age kids, if I could train them to do a job that I thought would be future Proof, that's, that might be it.
That would be the time job. com. org 300,000 strong community.
And I'm telling you, those are the two fastest growing elements in that community. Is data engineers and data and data scientists. Yeah.
Yeah. Yeah. Because that's where the action is, at least right now anyway.
And going forward, it's Interesting. Oh, it's in the way we, you know, in the, in the late nineties, we, the network engineers, right? Remember, remember how important those roles were.
They still, obviously, it's incredibly important. That's how I got into computers playing with Word Perfect. And Novell.
Novell, You're doing your IPX land. Yeah. Remember I was, I don't ask.
Anyway, Evan, it's been great catching up with you. Let, let's, we gotta give some call to action. You know what, okay, a good, a good marketing person I had here earlier said, you always gotta end with a call to action.
So the call to action is, is if folks are interested in anything, any resource associated with building our role is to make it super easy for developers to start with Time series. The database is super easy to use. com or do your local chat, GPT search and see how we're doing on go.
How are you doing on go? I think we're doing pretty well. I got a note today that a customer told one of our people, we found you.
We did a, a search on chat g PT for DevOps and SRE. Oh. And you guys came Up and they, and they sent them to, yeah.
That was pretty cool. Yeah. It's the world.
That world has changed. You could do a whole show on that. That's A whole We do anyway, my friend.
It's great seeing You. Good to see you. All right.
Thanks for having me. I think Kaplan, he's thanks and thanks, uh, AUSA folks for sponsoring this. Absolutely.
Well, as a matter of fact, we have more Susa coming up later today, so stay tuned for that.