PromptLayer Founder Jared Zoneraich on Building Reliable LLM Apps for Production Use
In this interview, PromptLayer founder Jared Zoneraich walks through a modern workflow for building production-grade LLM apps: version every prompt, run automated evals and cost checks, trace live traffic, and build regression sets that domain experts can tweak in a no-code editor. We’ll dissect real failure stories, tooling patterns, and team rituals that turn fragile prototypes into dependable AI agents powering customer-facing products.
Transcript
Hey everyone. Welcome back here to techron tv. My next guest is Jared Zona Reich.
Jared is, I think I got it right. Jared is the founder and CEO of a company called Prompt Layer, and it's his first time on Techstrong tv. So let's welcome him.
Hey, jar, how are you man? I'm great. Thanks for having me.
Pleasure. So, you know, you found, you're the founder, CEO here. Tell us, I mean, you didn't wake up one day I, and said, oh, I feel like doing this today.
What, what drove you, where's your passion for what's behind Prime Player? I feel like doing ai. Yeah.
Um, don't we all? Yeah. So we, we started, we, we launched this product about two and a half years ago.
Uh, me and my co-founder and my co-founder and I known each other for like a decade from hackathons back when we were, we were in high school and, uh, we, we were working on something before it, it didn't really have a lot of traction. And Chacha PT had just come out and we're like, how, what are we gonna build with ai? I was at a AI startup before this.
He was doing machine learning research. So we actually had a list of ideas and prompt layer as it exists today was the idea we made to get through those other ideas. So we said, let's make a tool to just help us build things.
We ended up releasing it in the January after Chad Beauty came out like a month later, and it blew up. And we said, all right, I guess we're prompt layer now. Go figure.
Sometimes just serendipity, right? If you build for yourself and you're customer number one, it makes a lot easier. Well, you know, so Jared, over the years, it's been a lot of years, I've spoken to a lot of founders and you know, that's actually a very common kind of pattern, which is, you know what?
This, I've got this problem and I don't see a solution out there for it. I'm gonna build the solution for my problem. And now, hey, I can't be the only one who has this problem, right?
There are other people who probably want this solution as well, and there's been some amazing companies founded on that very premise, so it's all good. Um, Jared, you, you talked a little bit in here about like kinda what drove you to, to, you know, found promptly or two and a half years ago, fill us in on the last two years. And, you know, I'm sure things have changed, right?
The the, the solution has changed, the problems people have, the use cases. Give us, give us the background. Yeah, totally.
So I think where we started, uh, so lemme lemme start. What was the first product? It was basically a tool.
We were building different hacks with AI and building different projects, and we wrote a prompt and then we updated it and we forgot the good prompt. So it was pretty much just a logging observability tool to remember that last good prompt. That's still a core part of what we do, but we do a lot more now, like you said, we're really, we like to consider ourselves the, the workbench, like the all-in-one workbench for building AI systems and building AI products and workflows and building them as a team.
So what that means is we, at the core, we're prompt management. So how do you version all your prompts? How do you log them, how do you edit them?
But also how do you eval them? How do you test them? How do you, how do you unit test these agents and make sure the outputs are good?
And then how do you log the results? So it's really this way. Teams iterate.
And I think our core thesis, which we've developed, honestly, we didn't start with it, we're both engineers. We built it for engineers, but slowly our customers kept coming in with non-engineers on the team. And the first time that happened, it was actually a, like a parroting coach and someone, they, they brought on someone who was a teacher for 15 years, she never coded in her life.
And we said, what are you doing? What, why are you using prompt layer? What's going on?
And she explained to us that the engineer set it up for her. She goes into prompt layer, she edits it and opens the app. And that's how they tailored the voice.
And she was the one who knew what the output should be. If you're building legal ai, engineers don't know what that is. Or, or psychology, mental health ai.
So that's our core thesis. Now, the best prompt engineers are not machine learning engineers, and you need to include the whole team in it, in the process. It's not just an engineering pursuit.
And it's really, it's a new type of knowledge work that, uh, is, is being built. I, I agree. You know, I I've been saying for a long time, you know, they estimate now, I don't know, there's 40 million something, maybe, maybe 40 million people around the world who are code, who code.
You wanna call 'em developers, call 'em developers. But with ai, we're gonna have 500 million people Yeah. Who code, right?
Right. Who, who, who develop stuff and, and that's gonna be the world, right? And so we need, we need prompt layer like applications that are gonna service these 500 million people who are developing apps, Right?
Right. There's a, there's a new, the brand new persona that's gonna be involved with building software, and it's not someone who has a PhD in machine learning. It's these new type of coders, like you're saying, it's the, uh, it, it, it's really the person who knows what the output should be.
And that's what makes it so much more powerful. And I think there's a lot of people, you know, there's a lot of opinions in AI these days and, uh, one of the opinions people take is the opposite that, hey, this is still machine learning. We need to get really deeply technical, but I think those teams are gonna lose because frankly, engineers, I'm an engineer, so I'm allowed to say it.
Engineers don't know what the output should be. A lot of times they're not the ones who have the sense and the taste and the product insight. They, they, right.
They don't have, they have the engineering background, you know, the coding background, but they don't have what the, what's the special source of a particular app or, or a subject matter expertise even, let's call it. But you know, Jared, in the low code, no code space, right? We called these people citizen developers.
Yeah. And a lot of people didn't like that name. A lot of people did like that name, but a lot of people didn't.
And you know, the idea is, is that they're not professional developers. They're, they are developing, most of them are developing just one app for a very specific purpose that suits their, you know, whether it's their hobby, their passion, their job, or what have you. They feel driven, I need an app that does this for me, not really looking to learn best practices or get a PhD, as you say, I just, this is what I want, right?
Yeah. And I, I, and I need it. I need it for work, or I need it for me or whatever.
Um, now look, in the last couple years we've seen, you know, crazy amount of evolution in terms of how people work with prompts, how people work with LLMs, how we were discussing it on Textron Gang earlier today. You know, when it first started, all these LLMs were based on, you know, uh, uh, a a a fact of, you know, it was a 2-year-old ball of, of, you know, a snapshot of whatever they were using a subset of the internet today, when you, you know, talk to your prompts, they're going out live on the net and grabbing stuff. So it, it, that makes the world very different, modern, in a modern workflow of, you know, building LLM coded apps and so forth.
What, what's changed? Is there a best practices, Jared, or are we still sort of in the evolving practices? Yeah, a little of both.
A little of both. I think, uh, I, uh, there's a few things to say here. I guess.
Firstly, if you're building an AI application or building something with lms, uh, perfect is the enemy of, uh, of, of, of, of duck. Oh, good, of good. Uh, the problem with perfect a lot.
I see so many times, so many teams, engineering teams, real hardcore teams, they're trying to get this perfect prompt and this perfect eval where they get a score, a metric that says how good it is, and they can fix it. Um, you need to prove that the AI application is gonna work to your team so you can invest more resources to build these evals. And a lot of times the way to make the prompt good is to start logging how people are using it.
Because the cool thing with LMS is there's so many different ways to use it. Uh, it's a new paradigm where users can enter anything basically into your AI application. So you need to start tracking what they're doing, using those for back testing and historical testing or regression testing.
And until you release it, you're never gonna really have that. So we really encourage teams start best p first, best practice, start simple, get something kind of working, and then you could build the edge cases from there. And don't start with like a a 50 node agent that's doing everything.
Start with a very simple stateless prompt, and you'll be surprised how good it gets. And then the other best practice is to do it in a clean and organized way. I think, uh, the, one of the big problems with these things, and what scares a lot of people is you have prompts scattered in your code.
You have strings everywhere you have, you don't know which version is which. What our tool focuses on is how do you organize these? How do you, what were GitHub basically for prompts?
How do you version, how do you know which one's working? How do you connect these with tests instead of, it's the wild west, uh, but a lot of stuff is up for grabs. How much can single prompts do?
How many times do you, how many models do you need to combine, which is open source model? Which model's best for this? And my answer to all of it is assume it's a black box and it's just try and check.
And that's the best way to do it. That's good advice, man. Good advice.
It's a good way to say nothing at all. It's, Well, no, but I, I, you know, I think, so look, I've been through a lot of tech cycles in my career and, you know, there are people who sit on the sidelines and kinda wr their hands, should I do that? Should I do this?
I I always believe you jump in and that's the best way to learn, right? Mm-hmm. And so that kind of fits my personality.
Um, if someone's out here and saying, you know, I like what Jared says. He seems to have a common sense approach to it. This prompt layer thing sounds pretty good.
I'd like to check it out. What, how would, what would you tell him? How do they go about doing that?
com. Uh, you could sign up for free. Uh, we have a lot of free individual users, and we have a lot of really like enterprise hardcore engineering teams using it.
So it's a pretty, uh, pla a pretty a platform for a lot of people. You can follow me on Twitter as well. Uh, we give some updates there.
And, uh, yeah, encourage you to sign up and let us know any feedback you have. I love it. com com.
Yes. Excuse me. COM Hey, Jared, I hope this, I told you it was gonna be an easy interview.
I hope this was okay for you thing we left out. You think any effect, anything else you want to share? Probably the, the thing I'd say is it's not too late.
We're still super, super early in AI and LLMs, and whether you don't know how to code at all, you're a content writer who doesn't wanna get left behind, or you are a machine learning engineer who wants to be the one who runs LLM systems for your company. It's not too late. It's very easy to figure out these things.
At the very core, it's a input and an output. Don't even think about how it works and, uh, you'll become an expert very quickly. I love it.
com here at Ontech Trunk tv. We'll take a break. We'll be back.