Diego Camizotti & Rodrigo Domingues – Shaping a Financial Service’s Cloud Strategy Using GitLab and Terraform
Join us to learn how we helped one of the largest financial services institutions in the world shape their cloud strategy using GitLab and Terraform. Starting on a cloud journey brings so many questions around resource provisioning & management, security, compliance, how to enable the team with easy access to definitions, and keep everyone updated. As we know, the most reliable source of truth is the code, so the use of infrastructure as code paired with an inner-source process is a solid foundation.
Transcript
I'm thrilled to introduce Rodrigo Domingues, a principal architect at CI&T and Diego Camizotti, a master engineer at CI&T. Their talk is called "Shaping a Financial Services Cloud Strategy using GitLab and Terraform". Is cloud native fit for Fortune 500 financial services companies?
How do you go about adopting the future while at the same time meeting all the stringent requirements? Rodrigo and Diego take us through their experience doing exactly this and share their key learnings. Let's hop in and check it out.
And don't forget, you can ask questions and engage in conversation in the chat during and after the time. Hello, everyone, thanks for being here. We're going to talk a little bit today about this journey that we went through with this big financial company throughout using TerraForm to create their cloud strategy.
But before anything, let's start by introducing myself and my friend here. So I am Rodrigo Domingues. I am a principal architect at the CI&T.
I've been working with the technology for the past 15 years and I'm currently in the Bay Area. I love several different stuff, but coding is something that I'm truly passionate about. I love also camping, going off road and my guitars.
You can see them here. They are my babies. I am full-on old school Java Developer; Java will always be my heart.
That's what started my passion for coding. If you want to get in touch please feel free. Those are the ways you can get in touch with me.
And we also have Diego here. Diego, can you just say hi? Hi guys.
I'm Diego Camizotti. I'm a master engineer at CI&T and have been living in the Bay Area for the past year. Also have been doing lots of cloud and devops works and very proud of it.
I love coding, playing video games, playing guitars as well, riding motorcycles and lots of other things. And as opposed to Rodrigo, I don't have a favorite language; I favor which whichever gets the job done easier. And please feel free to get in touch with me, these are the ways that you can do so.
So let's jump right in. Yeah, but just really quick, we both said we work for CI&T; CI&T is a digital solution provider. We partner with several clients.
Most of our clients are Fortune 500 companies, big companies. We normally work with them for several years. It's a long term relationship both with them and with the people inside CI&T.
And we partner up with a bunch of cool technology guys, including GitLab. So if you need our help, please feel free to go in your browser and look at CI&T website. OK, that's it.
Well, let's jump then. What was our mission? Right, what we needed to do.
Let's start with that. I think that's a good starting point. What when there's a big client had some issues, what they wanted and first things first, what is this company?
Unfortunately, I cannot give you guys any names, but what can I say? So you have an idea of what the company is and what the challenges are. They are a Fortune 500 company.
They are big, big financial company. They are distributed worldwide and they have more than 60 people working with them. So you can imagine as all big companies, they have a lot of challenges and they have a lot of things that they need to work on.
And I'm just giving you guys these numbers so we can have in the back of your head like the size of the company and all the complexity involved in those elements. Right. And also, I wanted to highlight that financial companies normally are not seen as someone who can do cool stuff on the cloud.
And we are here to show that they are doing really cool stuff on the cloud and they're really, really working in good new ways and an approval that is this specific company. They already have been working with two previous cloud providers. They already had AWS and Azure on their portfolio, but they wanted to increase the diversity on their cloud elements.
So they wanted to add Google cloud. And that came from their own teams who started using Google Cloud internally for some pieces here and there. They have some projects shared across.
And he said, well, it seems to be getting traction. So how do we make Google Cloud a reality for the whole company? How do we bring all those security elements, how we can make it secure for so everybody can leverage the core products that Google have?
So it started from the other guys in the company and projects being created and now they want to OK, let's try to bring that to our offer lifecycle as a whole. And they didn't want to just add a new cloud provider the way they were working with the cloud at that point. They had a bunch of processes and procedures and all the compliance elements that they needed to make sure that they are working to secure way in the cloud.
But not all of it was automated and was not equally automated throughout the cloud. So what they decided to do on this specific project is, OK, let's create a new way of working with the cloud, a way where we can ensure all of those elements from a more automated way, where we can create all the resources in an automated way. So we don't depend on specific people having access to elements of the cloud and doing manual configuration.
We both know that humans are good and doing. But repetitive stuff is not our forte. So normally there's problems with that, so there's a little bit of the background.
And I want to dive in a little bit more on the challenges specifically that we faced. And I'm going to divide the conversation here in two points. The first one is the what I call the hard requirements, and that is, well, for that specific cloud provider, we needed to create the DNS structure.
We needed to create the network structure. We needed to connect the cloud with the on premise of all of those requirements and all of those aspects, the hard aspects of creating the infrastructure. And now we're not going to touch too much in here just because it's really specific to that company.
If you are on that journey, you're going to have your specific requirements. That's going to be driven by your stakeholders, by your needs. So it's going to differ a lot.
But one thing that I think we can share with you guys and it can be used throughout are the requirements that they wanted to achieve or the outcomes that they wanted with this project besides the actual infrastructure being set up. And those were, first of all, security rights. And I think everybody thought the same thing when I said financial institution.
So the first thing that pops to the head is security, security, security. And in this specific case, they had two main concerns in mind. The first one was data exfiltration, which just means they didn't want their data leaving their infrastructure.
It can go from premise to cloud, but it cannot from the cloud to the web. So how do I prevent that? How do I create an environment that when I can let people use what they want, but I have that security.
And the other piece is, they didn't want to restrict too much what people are using. They want to actually go into the other direction, which for some people here may say what a financial institution wanted to go that way. But yeah, it was a really, really good and refreshing element to hear from them.
They wanted to allow their engineers to do whatever they wanted on the cloud, but they wanted to monitor everything because if something happened, they needed to act quickly. The other piece, of course, is compliance. Right of big financial company has a bunch of rules they wanted to use and they needed to use.
And like I said before, they wanted to allow for anyone to use all the resources and any resource that the cloud provides so they can actually leverage the speed of the platform going forward. But they wanted to do that in a standard way. So I want to make sure that when Diego asked for a resource, that resource that he's using on the cloud is actually following the standards that the company has.
And it doesn't is a lot to do a bunch of different elements with it so you can use it. But there are rules around it. The other piece here is they wanted to have one source of truth, right?
We know with big companies that different areas on different pieces, Social Security, infrastructure operations, information security, that's different from network security. So on all of those pieces, they need to come together. And they wanted to have one place where they can go to and find what the requirements are and how they implement that.
They didn't want to have different sets of three different blog posts, different documentations that people need to ask, permissions from different sources. So they need to manage all of those, ask if they want to, OK, how we maximize all the pieces in one place so people can have just one place to go and discover what's happening. They also wanted to make it easy for teams to onboard and contribute to the platform, and what I mean is if you don't know anything about Google Cloud all if you don't know anything about infrastructure automation, how do we start the project?
You don't need to have there are teams that didn't need to have experts on the technology to be able to start. They didn't want that barrier to allow more teams to adopt technologies. But while also what they wanted is OK, if allowing this all the teams to use, we do have some experts around the company.
We actually have a bunch of them. So how do we make sure that those guys who are experts and can contribute back to the platform, evolving my platform as a whole, evolving my configurations on security, on compliances or even or resource optimization? How can I accept those contributions?
I can validate those contributions to relative how can I allow them to contribute? So a little bit of two requirements, but it was good things that they wanted to do. And also on the same level, what they wanted was they want to allow for a small incremental changes and also corporation level changes.
And normally the flow that they want to follow is the same that we follow for programming, for an application as a whole. You want to deploy a small change in the development environment once that's OK, you go to the secure environment, then you go to the UET environment and you go to production. That's going to flow with the infrastructure.
Same thing. They wanted to test it out progressively, but also because it's infrastructure. They have the requirements of what happens if we discover an issue here.
What happens if we discover that we have a security hole we need to fix right now? So they wanted the ability to do it incremental, but also, if required, X in one moment and apply something to the entire environment. So that was one big piece of requirements for them.
And finally, what they want to do is to allow cost management. And I don't mean how cheap or how low price it can be. Something to run something on the cloud is actually when you have a company as big as they have, you have several thousand projects going on.
So how do you make sure that the area who requested that project is being charged back in the correct way, in a way that it's been built by what the resources they're using, not getting resources from other companies? How do you share what do you share the cost of the shared environment? So they want to allow for cost management.
They want to be able to discover what's happening there. And, you know, once when me and Diego had heard all of that, as you can see, there are requirements that contradict each other. There are a bunch of complex thing to do.
But it's not easy peasy. That's the easiest thing we can do. Right?
Just give me a cup of coffee or a mug of beer Sunday evening, and I can do that. You know what I said to say? Hey, Diego, can you come up with the solution?
And yeah, and that's what that's when I came and said, so you want to build something that's secure, that's a company wide, that's sixteen K employees worldwide, allow them to do everything they wanted. But in a secure and compliant environment? Sure.
Not a problem. Just open your favorite text editor like this and go do something else. Could come back a couple hours later and it's done right.
That would be magic. But that's not life. And life is a bit different.
So to solve this, we thought of three big pieces of this puzzle that could help us. And whenever we're talking about something that's multi cloud, that's called as infrastructure, as a code and code as source of truth, nothing comes to mind more than Terraform. So Terraform is easily integrated with any any cloud provider out there.
It's easy to start and on-board new teams once everything is set up and allowed us to do everything that we needed to do on the cloud providers. So the first thing that came to mind when we started structuring Terraform was modules. So how would we structure all of this that could allow us to do some very easy ways to roll out changes throughout lots of codes and things like that.
So these are one of the key modules that we thought that could solve things for us. And then it was really easy to start spinning up new resources based on that, because then we just created a new element on the project on your network or whatever that used the modules, defined the environment that it was going to run and bam, it was running that easy. And it also allowed us to do some very easy rollout changes throughout the whole infrastructure, because then we could simply changing one point and just reapply everything that depended on that point.
And simply everything was fixed and working as it should again. And also it provided a very easy and reliable way for us to share the information. So if your project on Google Cloud, in this case, neeedd to use a network to create an environment, you don't need to hunt that name.
You don't need to find it hidden in the dark or something. You just call the Terraform module that built that network and say, hey, give me the name of that network that I need to use. Put that in your code and give it gets to go.
And the safe part is, the only people that would have access to this are the people that actually need access to this because they would be behind a permission firewall, inside the bucket, inside Google Cloud. And talking about the second big piece, it's GitLab and not just as a source code management, but also as a CI tool, since it would provide a very the flexibility that we need to connect to several different cloud providers to do everything that we wanted to do and also to allow other users to contribute easily, since it can provide in their sourcing through merger requests or whatever pipeline was to find. And I'm not a big fan of monorepo strategy, but it made really sense for this project that we were working because there were several things that were tied in together.
And if we would have separate repositories for each of these things, it would be a nightmare to manage everything. So it made lots of sense to go with one big repro that had everything and also would be very easy for us to find stuff and to have other people also contribute with what we were doing and what they would leverage in the future. And then we started talking about the CI part.
How would we do so? The key point of GitLab CI here is to be the point of control of what happens and when it happens. So whenever a developer did get push force and please just use force, if you are a Jedi master, then GitLab would be able to identify what needed to run where, what was being changed.
If it's just a validation, if I need to start applying things and do things. And the cloud provider here entered as the cloud provider here and served as a place where these jobs would run on this case, it was Google cloud provider using cloud run. It's a product that we used on lots of basically run things.
It also allowed us to keep configurations on get them focused on the pipelines itself. So this way, we didn't have to manage more than one different type of image for GitLab to run for things like that. We just needed one image that could connect to the cloud provider via an http call and do the calls for us.
This allowed us to also keep the configuration very focused on what needed to happen. So we had one huge GitLab CI file configuring several different pipelines that knew what to do and what stage and and what to call and what environment they want to run and things like that. So it allowed us to move very quickly as well on that.
And also in the end, it was only one service account that we needed to manage on GCP side. So it kept things very tight as well. And the last piece that we had here, the very important piece is the cloud provider on this case.
It was Google cloud platform, but it could be any other of these. And it had only one responsibility of this lifecycle and it was to run, run, whatever it is that we needed to run, be it a Terraform validate, Terraform apply, be it a resource python code or whatever it is. Obviously, it's much more than that.
But in essence, that's what it did. And doing a quick recap against the challenges that we had just to make sure that we were able to capture everything Terraform and GitLab has have pretty similar contributions to all of this, but because they allowed very similar things in conjunction. Right.
So source of truth was in the Terraform, but easy adoption were both; compliance also were both, because Terraform has the project structure, how things should look like and GitLab had compliance validations to check that everything was made right. Also with Terraform, the way that we were able to structure it made it very easy to do incremental or both changes, whatever we did it to do. GitLab provided us the flexibility to control all of these executions whenever we needed it.
So it was very easy to have these two working together. And finally, the cloud provider provided all the resources of the security and monitor that were configured through Terraform code, provided all the cost management and preventing data leak. But again, all this was built based on what was defined on Terraform, executed through GitLab and finally actualized through the cloud provider.
And whew, that was long. So at the end, we had several key learnings that that we bear the scars and so today and we want to share a couple of them with you guys. So the first one of the obviously about Terraform is fix all uyour versions that that's not just to have Terraform version itself, but the versions of the modules that you're using, not the modules that we created.
But behind that, we were to use the other modules that are either on GitHub or GitLab or throughout the Internet. And to thing that the key message here is make sure that you trust where did that module is coming from, because not everyone has the same strategy in version tagging or something like that. At one point we had something that stopped working, simply stopped working, and we didn't do any change to that module.
And when we realized the module that we were using from GitHub had have been updated on the tag. So someone updated the tag and that introduced a bug that something for us. And we spent a whole day trying to find what it was and fix it.
So at that point, we decided that for some key modules that that happened quite a bit, we just made up fork and used the fork versions that of the live one. I know that's not ideal, but when you have something practical like this, it's good to be sure of what you are doing. And also be aware because Terraform in the end can execute anything.
So if there is like python code underneath that you're not aware of, you might jeopardize everything that you're building because of some malicious or not throughout testing code. And these are the messages about terraform, technology and Rodrigo and also has a couple of scars to share. And I want to share with you guys to make things right.
The first one, unfortunately, terraform is not Java. So, yeah, I know. I said especially for me.
But the thing is, what I'm trying to say to you that it's not coding infrastructure is not as the same thing as coding application. And yet we've both been on this journey of infrastructure and creation through code to a bunch more time. But when we started we were application developers and we still develop applications.
But when you compare, those things are very different worlds. Right. So when you're thinking about validation, when you're thinking about testing, sequencing of things can be tricky.
It's different than what it is on an environmental application. So you need to be aware of those things the other basis if you want to go to the infrastructure as code, please, please do not go without a CI CD pipeline. Right.
Like I said, infrastructure has code has its own challenges. First thing, everything is production. You may look at the environment.
It's being called dev, but that dev environment that dev network is used by all the dev environments. So if you break the dev network, you break the entire environmental so that's production. Right.
You need to think as every environment or ever changes a production change. So you can need to keep that in mind when when designing your process and all that. And you will do it.
But then if you don't have a CI CD to help you make sure that you execute the same steps every time, you can break stuff. And it's not funny, right? There are things that cannot be rolled back.
If you go to cloud environments. There are some things that even if you roll back, the resource is not released immediately. For instance, project names in the Google cloud platform.
Even if you if you create a project, even if you delete it from in order for you to use that project, then you have to wait at least 30 days because that's decreasing period or actually deleting the project inside GCP. So you need to be aware of those things. It's not a huge beast, is not a complex thing that is going to completely change your work, but you need to be aware of those differences.
Those things will give you a lot of headache if you don't think of them from the start. But again, go with your open heart and you'll be happy. It's been a great journey.
And that's it. Everyone, thank you very much for your time. Thanks for paying attention a little bit.
Please feel free to reach out to any one of us. These are our contacts. Again, we're happy to discuss any elements around infrastructure as code, CI CD, GitLab, please just reach out and thank you.
Thank you very much, everyone, and stay safe.