Exploring ScyllaDB with Dor Laor at AWS re:Invent 2024
Dor Laor, co-founder of ScyllaDB, discusses the database’s development, its features, and the challenges associated with open source licensing. Key points include the significance of data management and the evolution of database technologies, with ScyllaDB highlighted as a resilient and fast distributed database solution.
Transcript
This is Techstrong tv. Hey everyone. Good morning.
It's Alan Shimmel here at our Techstrong Wind Studio, where we're doing our coverage of, uh, AWS Reinvent. You know, I got out here on Sunday, today is Thursday. I've had about enough reinvent.
It's been a great show though, along with, you know, me and 60,000 other people. There's been a lot of keynotes, a lot of announcements, a lot of catching up with industry friends. Speaking of industry friends that want to introduce you to do Lavo.
Do, did I say it right? Lau It's, uh, hard to get. Yeah, it's not so hard.
Do Ur um, do, if you don't know, is the Found or co-founder, CEO of a company called Cilla db, makers of the Cilla DB database. You may or may not be familiar with it, but hopefully by the end of this interview, you will be Do welcome to Text on tv. It's great to have you back on.
Thank you. Good morning. Good morning.
Wonderful to, to be here. Absolutely. So, Dore, before we get into Siller and News here at Reinvent, let's talk a little about you, right?
You're, you're a co-founder, CEO, but give us kind of your journey. Mm-hmm. Absolutely.
You know, as, as a co-founder and A-C-E-O-I get, I feel married and kind of merged the company, but, uh, you don't Have to tell me. Yeah. I do have life and I mountain bike and ski.
So not only that, and family, of course. Uh, if we watch, um, so, um, I'm, or, uh, originally from Israel, eh, I co-founded RA 12, uh, years ago, together with my partner Avi. Eh, we're strong in tech, in low level computing by my, my first, uh, company I was involved with, uh, as an employee was, uh, Charlotte Web Networks, uh, interesting name.
We, we tried to compete with Cisco's core business. They bit router, and we created one, and it worked, uh, later on with Avi, uh, um, and Benny, my chairman, we created the KVM hypervisor at another startup, uh, eventually acquired by Red Hat and KVM runs, uh, AWS, uh, cloud and Google Cloud. Yeah, no, it's one of the most popular ones out there.
I remember those days. Yeah. It was, was some win.
And, uh, we used to fight against open source, Zen n VMware, and, uh, eventually KVM, like they became a standard in, in cloud computing. Mm-hmm. A wonderful ride.
And afterwards, we, we came up with the C db initially, uh, we tried to shoot at the operating system domain and take Linux down. Linux is so strong we couldn't take it down. We, we, we had an operating system for virtualized environment.
You're not the only one to, to die at that mountain. Yeah. You know what the funny people have tried to get.
Yeah. But It's, it was a worth worthy ride. Absolutely.
The, that was, we created still exists today in, in open source and still kicks ass Sorry if I Five. Okay. So it's an adult channel.
Um, eh, and then now we switched to the database, uh, the development, which is fascinating. It's a big world Out there. Absolutely.
And with the increased focus on data, as we've seen here at Arena Event mm-hmm. You know, data, it's about the data. You know what's interesting though?
You spoke about hypervisor. Here we are. So, I, you know, I first became aware of the whole hypervisor world.
I, I guess it was early 2000, maybe 2001. And, um, I, you know, VMware was part of spun out of EMC Dell the whole, but it was 25 years ago. Mm-hmm.
Right. If you go to the opening keynote here, uh, Tuesday morning mm-hmm. They're still talking about migrating from VMware mm-hmm.
Onto an Amazon stack or, you know, and then as popular as Kubernetes. And the whole cloud native thing is still hypervisor. Mm-hmm.
Right. The hyper, the, the hypervisor doesn't go away. And it's the same thing with the Linux now.
No, we haven't replaced Linux os, but if you look, there's been a lot Rocky Linux, you know, 'cause Red Hat's done some things with their Linux. And so other people have said to lessons, other people have come out. So though it may look like a monolith, there are cracks, there are changes that you see happening.
Same thing with database. Right. Another announcement here this week by Amazon, you know, they're claiming, uh, I guess it's based on Postgres, right?
Yeah. Or Aurora, the, the better, A better Aurora. Mm-hmm.
Uh, just, uh, folks in my team said, oh, we need to add a Postgres, uh, compatibility and SEL ourselves. I told them like, look, I definitely want to do that, but let's be focused. Uh, and I gave Aurora as an example.
The, the, the great team at AWS have been working for years and years with like huge teams and huge scales on Aurora, just to add it. And they compete with, uh, the plenty of other fantastic players like a cockroach and Hugo Bite and various flavors of, uh, um, of Postgres. It's, it's kind of, uh, you've got to be focused, and we're trying to be as much focused as what we do.
You know what, so I wasn't always on the media side. I've done tech startups most of my life, and that's, that's really the job of A CEO sometimes is you gotta know when to say no. Right?
You can't, you can't boil the ocean. You can't be everything. You gotta pick your fights, right.
Pick your bullets. And they, um, yes, there's a lot of Postgres kind of noise and tumble out here, but you don't, if it's not your, you know, it's gotta be in your priorities and you can't do everything on every time. Now, cila people who, what, what, what, what, what's, what's the special source?
Why do people want to, why would people want to use Cilla? Um, sure. So cila is a distributor database.
Uh, it's no SQL database. It doesn't, uh, support SQL, but the trade off is that it's, uh, extremely fast. It has distributed, uh, very res resilient for, uh, any type of, uh, high availability failure, disaster recovery failure.
We run a deployments across, uh, eight regions in, in one cluster. And all of them were zone aware. So extremely, extremely resilient.
NSS so three nodes can do a million operations per second. Um, so it, it's extremely fast. Or originally we rewrote Cassandra from scratch.
We, we started their project by stumbling on Cassandra. We figured Cassandra is a very interesting project, Cassandra itself, uh, to try to implement an open source alternative to Dynamo DB and Google big tables. Yeah.
And, uh, they have, and it's a viable, uh, project. The, the problem with Cassandra, it's more of a high level implementation in Java, not the best choice for, uh, low level software. And throughout all of our careers, we were working in really low level, like, uh, AVI checks, uh, every c plus last line to check how the assembly looks like.
It also, we, we check to maximize the, the bottleneck of every compute, uh, instance, network instance, uh, the storage at the NVME drive. So we squeeze every bit of, of the hardware in order to have the best, uh, throughput, uh, latency and efficiency. That's excellent.
That was an excellent, uh, description for the people. Um, for people who want to maybe give this a whirl, give it a try, find out more, what's the best sort of OnRamp? Mm-hmm.
So there, there are multiple options. Uh, one can just download the, the, the free open source version and another, uh, and running in Docker or run an enterprise trial. And we have a database as a service.
Uh, two thirds of our revenue comes from the database as a service. Uh, we can run in our account in the customer's account. Uh, so it's relevantly simple to, to consume.
There's free trial, free tier over there too. And so pretty simple to run. It really is simple.
And you know, like you said, there's a lot going on within the whole database space. The NoSQL look, no, I, NoSQL databases probably came out. They really hit their heyday around 2000, or, you know, we first became really, uh, aware of them 2008, 2009, maybe 2010 in that area.
Couch base, Mongo. Actually, it wasn't, there was couch and then there was, I forgot what base was they, they merged a big couch base. Right.
And, um, so it's a, a relatively mature technology at this point. Some people say, well, what, what's new under the sun? What, how else, how else do you address this?
Mm-hmm. But you guys announced some news here, or relatively recently for, for reinvent. Why don't you, if you don't mind, do share a little bit of that with us.
Absolutely. Um, so, um, for years we were trying to simplify what Sila is by, by having a one sentence of, uh, the power of Cassandra. 'cause we were, uh, we, we give everything Cassandra can can do with the wire compatibility at the speed of Redis.
Uh, Redis is super fast in memory database, uh, more mostly used for caching. And we can, uh, provide almost the same, uh, latency and in memory. Um, cache can give you, but from the disc, uh, with persistency.
So we used this, uh, uh, power of Cassandra at the speed of reds. The missing, um, piece was the usability of DynamoDB. 'cause, 'cause DynamoDB, uh, it preceded MongoDB.
The, uh, they develop, it's, it's, uh, internal development from, uh, 2004 only later became public. And, uh, like a lot of things that AWS do, it's relatively easy to use the, uh, as a service nature, uh, it's really easy to spin and it's, uh, um, the elasticity to add and remove. It's, it's more of a serverless nature.
You don't see the servers in, in the front and, uh, probably the, the rest of the industry, how they develop. It's more of a, uh, okay, let's take these servers and, uh, make them available to the users and also have a, as a service on top of that. But, but these, uh, it's hard to get away from those servers component.
Uh, what we've done recently, we, we did the major architecture change and we, we changed all of how we deal with metadata. We added consistency layer based on the rough consensus protocol. Ver very important for consistency and ease of operations.
That, that one big piece. It's, it's a big project that took us four years to roll out in, in stages. Uh, and the last bit is, uh, a portion called tablets.
The, the idea is to not to think about the server as in, in the data, the database cluster as big tables. They're gigantic tables. Uh, but divide them not just for like charted per server, but chart them into tiny pieces called tablets of, uh, five gigabyte, uh, each.
And those tablets are very elastic. So let's say if I need to load balance, uh, in my deployment, and I need to add double the amount of servers I've got, so we add servers, but, uh, then we need to stream the data to those servers. And normally if every server has, let's say, 10 terabytes, a relatively large amount, need to divide it and send half away, it's, it's a lot of, uh, streaming to do to be done.
And it takes time to do that. And until half a, their right in, in the past release, uh, would stream to the other server, then you would sit and wait and you use the old capacity. And this process, uh, could have taken, uh, hours, sometimes more depending, depend on the schema.
With tablets, it's all five gigabyte pieces. So five gigabytes, it take us about, uh, a second and two to send each piece. And immediately that piece become functional.
'cause all of the metadata is consistent, and the clients become aware, the new data moved from one server to the other. Uh, the clients do not need to know. They, they get notifications and it's, it's, uh, the streaming is becomes automatic, uh, super easy to scale up and down.
We do it best better than DynamoDB, which led the industry. So we can double, uh, your capacity in something like 10 minutes. It's a function of, of the amount of data, but easily double a big cluster in, in 10 minutes.
It's nothing and shrink back. Um, and if you're a customer, that's changes the way, uh, you, you do your planning sizing, uh, with Sila, because let's say many customers have the workloads a baseline of let's say 100,000 operation per second. And here and there, they may have spikes of usage for, um, like live, we watch live, the audience live up to 400,000.
Sure. So instead of be provisioned all year long with 400,000, uh, operation per second, it's more expensive. You get provision a base, and if there is a spike in minutes, uh, the, the database as a service adjust to the spikes.
That's fantastic. I mean that really, it's almost from what you're describing, I think almost like nanoparticles, but they like transformers, right? They transform into, you know, into one, one thing.
Um, this tablets is available in the database as a service. Is it also available in the open source, or that's strictly a, there's a premium kind of feature. It's also available in, in open source.
It is. There's differentiation, uh, with the enterprise, uh, visit. It's a bit faster, but, but the open source one is available too.
That's fantastic. Let's talk a minute, if you don't mind. You know, a lot of companies over the last year or so, we've seen a lot of companies changing their open source licensing.
Uh, you know, been a lot going on with how do we make a profitable business with an open source model. And I thought we solved that before, but evidently it, it's come back up now, you know, you're this co you're the CEO co-founder. You maintain a very robust open source community around cdb.
What, you know, how do you view this whole licensing? And are other people using it to create a commercial product and, you know, kind of living off of your hard work? How, how do you view all that?
It, it's definitely challenging and it's pretty much never solved. 'cause, um, e even companies like, uh, you mentioned Rocky Linux and, uh, uh, red Hat after SE, even a mature company like Red Hat need to constantly adjust the, uh, uh, what, what they do for free, what, what they do for paid. And, uh, uh, when I was a Red Hat employee back in 2008 to 12, there were lots of internal discussions about, uh, how not to expose, how, uh, um, red Hat Enterprise is being, uh, built and, and all of the composition of the packages.
But because, for good reasons, because other people, uh, clone it and come up with a free ride. Yeah. From understandable reasons.
But, uh, it, it, it's really tricky, this model. Yeah. Um, it's really, it's all open source, uh, work really well when, when it's not your core business.
Uh, so let, let's say if, if you're a Facebook and you publish your ai, then it's, it's not still your, it's supportive to your business, but not core. That's great for companies that the open source is core for their business. Uh, it's a give and take relationship.
Uh, and it needs to be adjusted. Uh, and we see what other database companies have done. Um, with, with licensing.
It's, uh, it, it's hard out there, uh, especially now, uh, where the, the entire market is, becomes more healthy. Uh, we don't have like, huge amount of piles of money, like 2021. No.
And the industry tries to be profitable. Uh, it pushes a lot of users to, uh, free users free usage, and that's creates a pressures on vendors. And it's more of a cycle.
So sometimes it moves here and it moves back and need constantly to figure out what's the right thing to do. And we adjust from time to time what we try to do to minimize the amount of change, uh, that what's available for free, what's available for paid. Uh, we, we do try to minimize and not to make a wave, uh, to absorb those waves.
Sure, Sure. Have we mentioned the website? Uh, Did we give the URL?
I don't think we did. Uh, so thank you. com.
And that's both for the open source and for the, uh, databases of service, or you get everything from one site. Mm-hmm. That's fantastic.
Do, thanks for stopping up. I hope you've enjoyed AWS reinvent this week. It's been a good show for you.
Yeah, it was super busy as always. Yeah. Uh, lots of meeting, like it's an industry show Together.
Yeah. Wonderful to meet everybody. And now where the, the brand is no more known.
Like, people stop me and say, oh yeah, sure. Need to, like folks from Korea folks, folks from, uh, Well that's the internet all over the world, right? Mm-hmm.
I, it, I still get a kick outta that when we do our webinars and people log on and say hello from here and hello from there, from everywhere, you know, and I, it kind of blows my mind still, even though I'm not doing this, I don't know, 30 years. Um, anyway, continued success with Cilla db. Come back, keep us posted and we'll talk soon.
Absolutely. Thank you. Alright.
Cilla DB here at AWS reinvent. We'll be wrapping up our coverage at the end of today. Uh, it's been an exciting week, and if you're watching this live, all of our, uh, interviews and content from this week will be replayed on text drunk TV next week.
So if not, to worry, if you wanna rewatch this, whatever, you'll catch it next week on Textron tv. Until then, though, this Allen Shemel four Textron, thanks for watching. We'll be back in a little bit.