AI Training and Copyright: Legal Rulings and Ethical Concerns | TSG Ep. 869
A recent court ruling permits AI companies to utilize legally purchased books for training, igniting debates on copyright law and ethics. The need for regulatory measures in AI development is emphasized to prevent potential misuse. Additionally, AI agents are being incorporated into CRM systems at no extra cost, disrupting traditional market practices. Concerns regarding social engineering risks in AI models highlight the necessity for governance and security protocols.
Transcript
Hey, everyone. Be advised, everything you say do right and put on the web, can be used by an AI train. AI portraying.
You're watching gang. Hey everyone, it's Alan Shimel, and I am really happy be telling you I am recording this at Platform Con in Person Day in New York City. We'll be here all day.
We'll be streaming live later, so do check. Well, by the time you're watching this, it won't be live, but you'll be able to watch it on, on recording hopefully. But we are here all day talking to some great people.
It's a, it's a vibrant high energy crowd at Platform Con in New York City. So, uh, good stuff. But we're here for Textron Gang.
We've got some, as usual, a good and bunch of AI news and different things going on. A little cyber here, a little AI there, a little DevOps there. Let me introduce you quickly to our gang for today.
There, people you've met here before, so I'm not gonna do the whole dog and pony show, but we've got John Swartz, Fred Wilmont, and still in Denver, the dean Mike Ard. And hey gentlemen, how are you? Good.
Doing great, brother. Good. Heading To the airport right after this.
Good for you. I'm glad you'll be going home. I'm actually gonna be in New York all weekend, so looking forward to being home still home for me.
Um, guys, let's just jump right into it though. Uh, kind of a, i I don't want, if I, if I should use the term a landmark ruling. Is this the high watermark for AI court rulings?
Is this the, or is this just the beginning? But a, a court has ruled that, uh, anthropic using books and other material that train their AI is perfectly fried. Even though those books don't belong to them, they're not compensating the people whose, whose IP it is.
Um, we don't know if the books were obtained even paid for or just hiring were paid for. I don't know. Mike, what do you think?
Well, I'm gonna kick this to John really quickly, but yeah, it seems to be that the court is saying, as long as you bought the book and you didn't steal it, you can do whatever you damn well want with it. But John, Yeah, That, that, that's basically, that's basic base. That's basically just, so as Alan said, this is potentially a landmark, but there are some caveats and some howevers in the decision that leave the door open for other different interpretations.
So basically, just to recap, there's a federal judge in Northern California who ruled that anthropics use of, of legally purchased books to train. Claude does not violate federal copyright law. And he's basically saying the use of these of these books to train LLMs was Quin, essentially transformative and was a, and was, did not violate the fair use doctrine under copyright law.
So the key phrase, one of the key phrases here in decision there was a lot of hyperbole, was that authors complained. There are three authors who sued anthropic. Their complaint is no different than it would be if they complained that training school children to write well would result in an explosion of competing works.
So they're saying basically that if you legally purchase the content, you can use it for training purposes. And, um, the, the authors themselves there three authors, they had thought and interpreted what Anthropic was doing as large scale theft. But here's the caveat.
So the judge whose name is Elop, said that Anthropic may have broken the law when it's separately downloaded millions of pirate books, and it will face a separate trial in December over that issue. Um, he also said that the decision does not address whether the output of an AI model infringe copyrights, which is an issue in related cases. Um, one last thing I'm gonna mention is that the court documents show that Anthropic knew there was something up and they had had to change course.
The anthropic employees initially were concerned about the legality of using higher ed sites to act work. So they had brought in a former Google executive in charge of Google Books, which is a searchable library of digitized books, and they kind of changed course and legally purchased books for the content training. So in other words, if you buy it and legally you can possibly use it, but all channel worms there, there are all sorts of repercussions and all sorts of different interpretations that can go on moving forward.
This is far from settled. You know, John, I I get what you're saying about how the judge sort of came to this decision, but you know, that that example you said about giving children a book that they was bought legally to teach them how to write properly or in a style, I think is, is a great example. But there's another legal theory of play here that I think is, is germane to the AI use case, which is that of derivative works.
So it's, it wouldn't be okay to give a book to a child and then he in, in essence takes liberties in, in using the exact words plots, storylines in writing his own book, or it is more akin to sampling in music. Yeah, Right. And so I'd like to see where the court came down in terms of a derivative works type of argument that this ai, when we talk about training ai, what we're talking about is AI then regurgitating that information for someone else's IP or for someone else's use without adequately compensating the original author of, of that it.
And so I, I mean, I, I think this is a bad decision in reading some of the language about it being transformative and so forth. You know, I think this judge is an AI fan who's Yeah, you odd with the magic of it. And, and you know, I think if you asked him, he might say he's making, what would they call in the legal profession a common good argument that the, the common good of having AI help society outweighs the individual, uh, IP concerns.
Yeah. I don't know if I agree with that, but that's what it sounds like in reading, reading the article. So do I understand this correctly based on this logic?
Therefore, philanthropic just needs to go send a, a check to the various publishers for the cost of the books that pirated and they're good to go. That's my interpretation. Um, you know, rather, I mean, that would be the safe, prudent, uh, path for them to take.
Um, you know, it's, it is interesting though because they were also piring of evidently a ton of materials. So they paid for like a small fraction of what? Of what, Right?
I mean, go going to buy all the books that you're training a large LLM on is not exactly Cost prohibitive. Yeah. And that's anthropic doing it.
They're one of the, you know, open AI andro, these are, let's call them ethical l lms. We're gonna talk more about ethical LLMs later. What about all the unethical ones?
Mm-hmm. It will be interesting to see if this case is appealed and I'm li I'm sure it will be what the interpretation of another judge will be, because I think in this case, this judge clearly gave wide latitude to AI. And, um, I I I'm gonna find it really hard to believe if, if other judges have fall in line.
I, I think there's gonna be so many different interpretations of this. We, this is far from over, far from deciding. Oh Yeah.
I, so I mean, this are Northern county. This is, this is kids c will let, let me, let me, yeah, let me, let me paraphrase. George W.
Bush, right when he described Michael Dukakis, right? This is from Massa. This is a guy from Massachusetts, right?
The most liberal state there. This is a northern California judge. I, I think other courts and other judges might have Yes.
A very different opinion here. It Was a home court decision. How about that?
Our home court ruling our Home. Fair enough, right? That, yeah.
All right, let's wrap up this segment. Great story by the way, John. We're gonna come back and we've got a few other things to talk about, including agents for nothing.
I don't know chicks for free. I'm, I'm channeling Mark ler. You're watching Text on Gang.
Hey folks, we're back in. Yeah, to Alan's point, there's this company called Creo, I think that's how you say that. And they make a CRM and they may not be the most widely used CRM out there, but they launched a bunch of AI agents this week.
And interestingly enough, they is not charging extra for these AI agents. They're basically saying AI agents are gonna be a feature and not an add-on. And one of the issues we have been seeing out there is every vendor seems to show up with some sort of AI capability that everybody, that they want everybody to pay extra for.
And it seems like, well, maybe that's not gonna hold, because at some point somebody has to be the, uh, startup that's trying to take out the incumbent. And one way to do that is to not charge for AI agents. So Alan, is this the market at work or what?
You know, I, I think we've gotta look at this in context of the world. Mike. This is a company, they're not Salesforce, they're not Siebel, they're not, you know, I mean, they, they have a nice CRM, it's kind of a, you know, one of the scrappy competitors.
And you are going up against an 800 pound gorilla, excuse me, who's promising to, you know, have tens of thousands of agents available for every purpose under the sun. And so how do you counter that and this, and since Time Memorial, this small guy says, I got a great idea. Let's give it away for free and we'll make it up on volume.
I don't know how successful a strategy that is, but at some point, agents do become commoditized and do become part of the offering. And, and they will be free at at least a, a base set of them will. And maybe, you know, much like open source, you'll have sort of an open core of free agents, but higher, higher, uh, functioning ones will cost a little bit more money.
I don't know. It's what I think, I think part of the exercise here is that the margins on software are pretty good still. And I can hide the cost of the AI agent in those margins and still make a, a nice piece of change.
And I'm just gonna be curious. But I think that this is gonna be a trend across the board. I think every other ISV is gonna come to the same conclusion.
I think for enterprises, AI may still be expensive because I gotta get the GPUs, but I think the cost of those GPUs is gonna be buried into the license somewhere. Now, you know, will the core license go up sometime they account for that. Maybe they maintain their previous, I don't know, Alan, what, what are the margins on ISVs these days?
I seen you remember it was somewhere in the realm, 50%, Um, net, triple net you're talking about, right? Mm-hmm. And, and you know, and it depends with, with SaaS involved, Fred's probably better suited to talk about that than me.
He lives it. Um, but it, you know, there's certainly margin there to put it in, let's put it that way. So, um, we inevitably, look, this is going to happen.
It is just a little earlier than I thought it would, but it also remains to be seen, are these agents real? Are people going to use them? What's going on with them?
You know? And I don't mean agents in general, though, I might be talking about that as well, but certainly in terms of the ones that create here is, is rolling, are rolling out. I, You, I'm kind of looking forward to using AI agents because I'm, there are just, um, all kinds of tasks in my daily life that I just scratch my head about and go, why am I doing this?
I'll give you my pet peeve example. So, you know, you write something in Google Docs and or in, uh, word, and then, you know, you gotta load it up into WordPress. I'm like, isn't there an agent that could just do that?
You know, just kind of like take that little thing and just, you know, 'cause it, it probably sucks up about, you know, two or three minutes, but it's just annoying as hell. But so, so here's my pet peeve. Is that an AI agent or is it an API call that, I mean, didn't it, didn't we have that kind of functionality already?
I think there's a couple of interesting things about what these guys are trying to do here. The first of which is they say, we're gonna, uh, allow you to build, um, you know, build your own AI agents. And the second thing is, uh, we'll let you choose your olms.
Uh, that's a, that's a provocative thing. If you're doing that kind of work, you're not, you don't wanna be held accountable to, you know, which model has the highest risk. I have this custom model built for my things.
This to me seems to strike at the heart of people that are doing an awful lot of AI work and have built things internally and are discerning members and say, Hey, cool, I can implement what I have here already. You know? And it is, I think shimmy, as you said, it is provocative to say, we can enc encapsulate this in our SaaS budget, right?
That, that margin, you know, everybody that would, would be a funder founder or an investor in a SaaS project would say, if you're making less than 80% right, you're probably doing the wrong things. So there's definitely room for that. I think the interesting part here that I still feel is pretty unknown is what do those ballooning costs look like?
This is a really big gamut, uh, to, to run here. If you're gonna suggest you can use any model, well, you know, and build your own agents, that that could be very, very expensive. Uh, it could be super optimized by people that know what they're doing.
It, it's pretty cool to see. I like the democratization of it for sure. It'll be interesting to see.
But you know, Fred, to your point, you look at like Amazon Q from what I, what I remembered, you could plug your own LLL, you could plug, you know, different LLMs into that. Mike John, who, who is Einstein? Is that Salesforce or one of the other big guys that's Salesforce.
Einstein. I mean Salesforce, yes. Einstein also allows you to plug multiple LLMs into the back end of it.
So again, I think this is Credo trying to match up against the 800 pound gorilla that they compete against every day. So to Fred's point, I think, you know, are, are folks out there building AI agents that are just gonna wind up being free pieces of applications that I license? And, you know, it feels like we're about to have this classic bill versus buying conversation for, Uh, I think there's a whole suite of folks that in order for people to adopt their technology, the latest thing is to build a plugin for their agent.
This MCP, you know, that agent so on. Uh, but it remains to be seen if those are gonna be widely used, right? To Alan's point earlier, we kind of just exchanged a way to interacting with, you know, uh, words, what we used to do with APIs.
Is it better? Is it works? Does it make more sense?
Do you get the right answers? I mean, still to be determined, but I think everybody feels as though in the fleet of agents, I must have an agent that someone can interact with in my, you know, my technology stack. So vendors I think are gonna rush to continue to do this.
And there's lots of public, uh, public agents today, and there's lots of public NCP servers and all of the things that go with it. But, uh, is that a trend? Is that sort of top out?
Do we get back to, hey, those things are, uh, probably not adhering to most of the general principles of security anyway, right? Is it, is it better practice? Is it, do you get more out of it?
Like what's the value prop there? And I think those are the questions. We, it's fresh, fresh snow.
So I don't know if we're gonna see that for a little bit. You know, ano, another lesson I've learned in, in years in, in the tech world is just because you can, doesn't mean you should, and I'm not talking about agents being free. I'm talking about, you know, creating an agent for every, to scratch, every itch that you may have.
It, it may not make sense to do that. And so, you know, I, I think we're going to go through a period where, where we are going to try to scratch every itch, but eventually, you know, when the cost of of AI gets factored into these things, even at a B2C level, you are going to do it where it makes sense. Not, not for every thing under the sun.
And, and that's again, following a, a, a, you know, a pattern that I've seen over 30 plus years. Mm-hmm. And you know, the temptation a among some of these tech companies, when I talked to 'em and Cisco execs said this, that hey, they're the thought that like 75% of your time is wasted on, like 44% of that time is wasted on repetitive tasks.
Like what? Or 31% is in meetings. So they're looking at this scope of 75% of what workers do throughout their career can be replaced in some manner or form by a R agents.
And that's kind of where they're coming from. It is overkill and it will not go to that percentage for all of us, but that's how they think. I, I believe, Yeah, that's just, we go bathtub curb, right?
You spend, uh, the first, you know, 35% of your time trying to figure out how to mue the data, executes their stuff, and then, you know, the last 35% of your time or 40% of your time trying to figure out how to interpret it. I, I'm not sure right, that whole class of line of thinking makes sense in this case, but I think the zero to one problem gets an awful lot shorter for sure. Uh, so that you can figure out how to flatten that first part.
But, uh, you know, the latter part of determining whether or not that's valuable and why it matters, I mean, still question marks there. I would say one thing on behalf of the humble API doesn't tend to hallucinate. Yeah, no, for sure not.
But you know, we, we'll see how it comes out. Alright, let's take a break on that one. We're gonna come back and talk about our third, uh, third uh, uh, segment today.
It's a thin Lizzy song, jailbreak, uh, you're watching Text on Game Discover Techron Group, the epicenter of tech innovation. We are your go-to for reaching IT leaders and practitioners worldwide. Our secret impactful content that sparks awareness, engagement, and top quality leads with us.
You'll access editorial websites, streaming videos, virtual events, custom content analyst research, and more. Join our satisfied clients. Let's revolutionize your tech journey.
Contact us today and tell your story to the world in the most powerful way with Textron Group. Hey folks, we're back and we're talking about jailbreaks and yes, it is a Thin Lizzy song and we dearly miss those folks, but that's another story for another time. But, um, Fred, there's a thing called Crescendo and there's a lot of different, uh, novel efforts to kind of get LLMs to cough up information that they're not supposed to.
And the latest one, I think is called Echo something or other, but it seems like this is a, a cascading series of discoveries that people are coming up with. And some are saying that this is actually not a bug, but just a flat out structural, structural flaw of an LLM. What do you think?
Yeah, so it's interesting that we've now parlayed what this is, you know, called the echo chamber attacks, what we're talking about here specifically, but it's kind of ai, social engineering. So same problem we have today in humans right now, what we're doing is we're, you know, the crescendo story that Microsoft put out explains a little bit how you can think about starting with relatively innocent questions and sort of evolving that to, you know, a harmful prompt, right? And then next thing you know, you've got poison poisonous seed sort of influencing the outcomes.
And some of those types of outcomes, uh, really can't be seen as a bug, right? It's just, it's, it's not a, it's not a flaw in the reasoning process. It's a way to manipulate the reasoning because all the models that we love, uh, dearly, right?
Or inference model. So you can see things like, you know, all of open AI's models, right? And, uh, even, even, uh, Gemini's, you know, two five models are subject to this.
Plenty of others are also. So Neural trust came out with a methodology here that's very interesting. It's interesting because of the, uh, the level of sophistication required, but also it's interesting because of the number of models it actually affects.
So the vast majority of models, right, are affected by this at least 40%. And then some of the ones we, we use every day, all the open AM models, that's up to 90. But the concept here is harmful prompts.
You know, let me, let me ask you how to build a Malta cocktail. Sorry, can't do that. Okay, well, lemme tell you what I know a, uh, you know, a Maloff cocktail to look like, and then, yeah, let me seat that.
And then let me get some of the ways to, uh, uh, interpret my seed with other questions. And the next thing you know, I'm now invoking, you know, this context, that's probably not right. And then we've got some paths that go down the, you know, go down the wherewithal of which direction we're gonna go with the model.
And I get an answer, right, step by step how to build a malt tall cocktail. And, you know, this is not, uh, this is not sort of a, a challenge that I think anybody thought we would not come across. There's a lot of, you know, governance, uh, guardrails being put in the market today.
But the challenge if we have here is that, um, the pervasiveness of something like this to happen and the influence of things when we look at the open AI requirements, uh, for security, right, not listed in there, is context protection and disambiguating, you know, malicious from good. So there are some interesting new definitions we're probably gonna have to come up with to understand how to appropriately regulate this kind of behavior and to understand when it happens. So can I say it's not a bug, it's a feature.
What? Uh, It's classic. I mean, but you know, we, we spoke earlier today about, uh, ethical LLMs, ethical AI players.
Again, we, we could try to build in better guardrails that'll prevent you being able to take this ai, you know, to me, it's a kid to putting age verification on porn sites. Are we really keeping kids under 18 from, from entering these sites? Um, I, I don't know if, if there's a good answer for this.
Not, not with present technology. And, and that's, if you wanna say, well, you know, put the brakes on AI as a result. Good luck with that.
That's not happening. Th this, this is going to be one of these classical security things where we'll care about it when people b***h about it loud enough. But right now it's, it's, the b******g isn't loud enough.
So on the guardrails that we're coming up with worth the damn in the first place, and I mean, why, maybe it's not just worth the effort and we should just let people know that this can happen and, you know, bad people are gonna do bad things. I, I think there's a bunch of parts coming out to talk about this, right? Where, you know, whether it's, you know, Google Secure AI framework, it's, uh, o OSPs new, you know, AI top 10 LLM, uh, concerns, uh, cloud security alliances.
Uh, there's lots of people talking about what the potential implications are, but I, I, I think there's a really good point to be made here. That's great. However, the tools don't all exist to guardrail all of these things.
And maybe we should quantify what the level of exploitation capability is first, and then we're backwards. I think we're, we're, we're over pontificating about the implications without really understanding the consequences. I also think it may differ by use case.
I mean, imagine your financial services firm, you know, and then people are teasing out information about your clients. Well, that's a bigger issue than maybe, you know, some retailer somewhere. So, um, you know, will this maybe reduce the appetite for using LLMs in certain use cases where there is a lot of sensitivity around the data?
So it's theoretically possible. Well, I think there's a conversation we had about, so if, if that's your risk, right? And the Kara put out this thing this week that talks about the AI model risk index.
You'll notice we've talked about, uh, philanthropics Quad, uh, uh, you know, four oh and where they sit on the security index a while back. But the risk index, if you're a large financial right, or, uh, you are using something specifically to guard, I don't know, the, the Koch formula, right? Then maybe the risk here for you is that is fundamentally too dangerous for us.
We're gonna implement our own model that's customized very specific as guardrails around it and doesn't require you as the foundation models and, you know, a public cloud infrastructure or something like this. There's ways of route it. I think quantifying what that actual risk is to you, this is a good indices for, for folks to look at, is if this is possible and that information is critical to me, then I've gotta make different risk decisions maybe about which model I use, which service I use, and how I weaponize that in engine for my user's benefit.
Fair enough. So, So I would think that, you know, in a couple of months from now though, when folks who experience that risk are paying attention right now are gonna wake up one morning and there's gonna be some sort of, you know, shocking ease of news that, you know, will surprise them, and we will be for us, we'll be like, Hey, we told you so. But I'm, I guess I'm, I'm wondering how long that might take.
'cause I think the, the bad guys are getting really good at social engineering of LLM. So I think, you know, this may be popping up on, I don't know, the front page of the proverbial Wall Street journalists. They used to say, uh, 30 days from now, John, what's your bet?
Yes. When I heard software engineering and I heard ai, I was like, here, it's, it's real and it's happening and it's, we're gonna read about it pretty soon. And that, I mean, that's always like the recurring theme among these companies in their rush.
They, they know it's, they know it's happening. They know that the what the risk is involved and they're willing to take it. So yes, Mike, it's matter of time.
Alright, I, I think that's going to, uh, which sounds like we're a little outta steam here, but that probably means it's time to end the show, which is a good thing. I, I've got a busy day in front of me at Platform Con, and as I said, uh, it should all be available on any of the tech strong, uh, TV outlets, YouTube, OTT, et cetera. Also, we've got a full tech, strong TV lineup immediately following today's gang show.
But until then, Mike, safe travel's home to New York. John, Fred, great to see you both. We'll see you soon on another episode.
For now, on behalf of uh, Textron Gang, this is Alan Shimmel. We're out.



