How Large Language Models Work for Title Agents | Ep 92
Episode Summary
Mo Choumil deconstructs how large language models actually work under the hood—from training on trillion-word datasets to token-by-token statistical prediction. He explains why AI hallucinates, how to separate high-trust tasks from verify-always workflows, and which models (ChatGPT, Claude, Gemini, Llama) fit title professionals’ security constraints. Learn the mechanical reasons AI fails, how to leverage context windows for 200-page commercial searches, and why mastering prompt engineering is the only skill barrier separating early adopters from competitors still manually drafting emails.
About Mo Choumil
Mo Choumil is CEO of Alltech National Title and host of the Title Agents Podcast. He guides title insurance professionals through industry transformation, focusing on innovation, operational strategy, and leveraging emerging technology. Mo is known for translating complex technical concepts into actionable frameworks for agency owners, top producers, and operations leaders navigating the evolving landscape of real estate settlement services.
Key Takeaways
- Large language models are pattern-based prediction engines, not rule-based databases—they generate responses token-by-token using statistical probability, not retrieval.
- AI hallucinations are a core feature, not a bug: the same creative mechanism that drafts emails also invents fake court cases with complete confidence.
- Title professionals must separate tasks into high-trust buckets (email drafting, summarization) and verify-always buckets (legal citations, recording fees, regulatory deadlines).
- Claude’s massive context window can hold a 200-page commercial contract in memory simultaneously, cross-referencing clause 198 with definitions on page 3.
- Meta’s open-source Llama model allows agencies to run AI on internal servers, keeping sensitive NPI data compliant and off third-party cloud platforms.
- The competitive advantage window is open now in 2026: professionals using AI as a first-draft machine are doing the work of three people while competitors wait on the sidelines.
- Prompt engineering is the only required skill—clarity, context, and structure in plain English determine whether you get vague output or workflow-changing results.
Episode Chapters
| Time | Topic |
|---|---|
| 00:00 | Intro: The 4:45 PM Friday title search scenario |
| 02:15 | Unlearning software: rule-based vs pattern-based AI |
| 05:30 | The librarian metaphor: why LLMs don’t retrieve, they generate |
| 08:45 | Breaking down LLM: language, model, and parameters explained |
| 12:00 | Training phase vs inference: trillion-word datasets and volume knobs |
| 15:20 | Why AI hallucinates: the mechanics of invented legal citations |
| 18:10 | High-trust vs verify-always task buckets |
| 20:00 | Choosing your model: ChatGPT, Claude, Gemini, and Llama for title pros |
| 22:30 | Real workflows: email drafting, title summarization, prospect research |
| 25:00 | The urgency of adoption and the compounding competitive advantage |
Full Transcript
Show Full Transcript (4,629 words)
In a world where change is the only constant, Mo Shamil stands at the forefront, guiding title professionals to not just grow their businesses, but to master the art of innovation. With every episode, you're handed the keys to unlock unparalleled growth and stay ahead of the curve. Get ready for a transformative journey. Imagine this scenario for a second. It is 4.45 p.m.
on a Friday. Oh, the worst time for anything to happen. Right. The absolute worst. And you are staring down this massive, like, 200-page commercial title search.
I can feel the stress already. Yeah, because a major closing is supposed to happen first thing Monday morning. And you just know, somewhere in that giant mountain of paper is, like, a buried mechanics lien. Or some bizarre zoning exception. Exactly.
And normally, I mean, this means you are totally canceling your dinner plans, right? You're brewing a fresh pot of coffee and settling in for just hours of grueling line-by-line reading. Yep. The classic weekend ruiner. But today, in 2026, you take that entire 200-page PDF and you just feed it into an artificial intelligence tool.
And exactly 14 seconds later, it hands you a perfectly structured, plain English summary. Flagging the exact break in the chain of title on, like, page 142. Yeah. That's, like, absolute super power, right? It really does.
But here is the terrifying flip side of that. What happens when that exact same artificial intelligence confidently invents a completely fake court case? Or hallucinates a phantom recording fee? Right. And you accidentally put that hallucination into a binding legal document.
Yeah, that juxtaposition is basically the daily reality for professionals right now. It's wild. It really is. So you're operating in a landscape where these tools can genuinely do the work of three people in a fraction of the time. But only if you know what you're doing.
Exactly. Only if the person at the keyboard fundamentally understands the mechanics of the machine they are operating. Which most people don't, honestly. No, they don't. Most people have tinkered with ChatGPT by now, right?
They've asked it to write a funny poem or, you know, maybe draft a quick email. But playing around with a tool and actually understanding its architecture are two completely different things. Totally different. And bridging that exact gap is our entire mission for this deep dive today. We are taking a shortcut to understanding large language models.
And this is specifically engineered for you, the professional working in title insurance, escrow and real estate settlement. Yeah, we're throwing out all the dense technical jargon. We just want to uncover how these systems actually process information. The mechanical reasons why they sometimes fail so spectacularly. Yes.
And most importantly, how busy professionals can actually leverage them to build a massive competitive advantage. It's all about that advantage. OK, let's unpack this. Where do we even begin to understand a technology that feels so, well, so alien compared to the software we've used our entire careers? I think we begin by completely unlearning our definition of software.
Unlearning it. OK. Yeah. So, let's say this technology was purely transactional and entirely rule-based. Based.
Now, think of traditional software like a highly efficient but incredibly literal librarian. OK, a literal librarian. Right. So, you ask this librarian for a specific book on commercial real estate laws in Illinois. They walk down a specific aisle, find the exact book and hand it to you.
And if the book doesn't exist? The librarian just comes back empty-handed. Software's built those old systems using millions of rigid if-then statements. Like, if a user clicks this button, open that file. Exactly.
If an email contains the word Viagra, route it to the spam folder. I remember when we first started calling those spam filters and credit card fraud alerts AI. We threw the term around a lot. We did. But looking back, that was a very different breed of artificial intelligence, wasn't it?
It was really just highly complex human programming. It was. It absolutely was. But these new large language models represent a fundamental departure from that. How so?
Well, they are pattern-based, not rule-based. Nobody sat down and wrote a line of code instructing the computer on, you know, how to identify a break in the chain of title. Wow. Really? Yeah.
Nobody programmed rules for drafting a warm but professional follow-up email to a difficult real estate attorney either. So what did they do? Instead of writing rules, engineers just fed the system enormous, unfathomable amounts of human language. And they essentially commanded the computer to figure out the underlying patterns on its own. OK, wait, wait.
If it doesn't have rules, how does it actually know what to do? That's the million-dollar question. Right. I mean, if a programmer didn't write a specific instruction for how to summarize a commercial lease agreement, isn't the machine just randomly guessing? Not exactly guessing.
Because how does it produce a coherent legal summary without an instruction manual? If we connect this to the bigger picture, pattern recognition at a massive scale essentially becomes its own kind of operational logic. OK. It's all an operational logic. Yeah.
It is a completely different paradigm. Let's update that librarian metaphor we used earlier. Let's do it. Instead of a librarian who fetches a specific book, imagine a librarian who has read every single book in the building. OK.
Very well-read librarian. Right. And after reading them, this librarian just burns all the books. Burns them? Burns the books and just sits at an empty desk.
So when you ask a question, they aren't going to retrieve a document for you. Because there are no documents left. Exactly. They are writing you a brand-new original essay on the spot. And it's based entirely on the mathematical vibes and patterns they remember from reading all those destroyed books.
That is a fascinating way to look at it. So they aren't looking anything up at all. They are just generating it from scratch. Entirely from scratch. This actually leads us perfectly into breaking down what the acronym LLM actually stands for because those three words, large language model, they explain the entire mechanism.
They really do. Let's tackle language first. Because the native environment of this technology is not computer code, right? Right. It is human conversation.
And that alone completely removes the learning curve of traditional software. I mean, you don't need to learn Python or figure out a complex user interface with a hundred drop-down menus. No, not at all. You literally just type to it the way you would text a colleague sitting across the office. The interface is just plain English.
Plain English. Which brings us to the model part of the acronym. And this is where the biggest misconceptions lie, I think. Oh, definitely. An LLM is a highly complex mathematical map of how words, concepts, and ideas relate to one another.
So it's not a database. I cannot emphasize this enough. It is absolutely not a database. It is not a search engine. When you ask it a question about a title statute, it is not opening a little digital filing cabinet inside its code labeled real estate law.
It's not Googling it behind the scene. Exactly. It is calculating a mathematical map of relationships. And the engineers map these relationships using something called parameters. Okay, let's pause on that word.
Parameters. Yeah, it sounds very technical. It gets thrown around in tech articles all the time. But what does that actually mean for the person using the software? Think of a parameter as like a volume knob.
A knob controlling the connection between two concepts. Okay, a volume knob. I can picture that. A modern LLM has hundreds of billions of these microscopic volume knobs. So when you type the word title, the mathematical model instantly adjusts.
It really shifts the knobs. Right. The volume knob for the word insurance gets turned up to 10. The volume knob for search gets turned up to 9. And the volume knob for something random.
Like banana. It's completely muted. The AI is constantly calculating the mathematical gravity between words. Which brings us to the final letter, the large enlarged language model. And honestly, calling it large feels like the understatement of the century.
It really does. Because to get those billions of volume knobs tuned correctly, the AI has to consume an ocean of data. A literal ocean. We are talking about models trained on approximately one trillion words. A scale that is genuinely difficult for the human brain to even visualize?
I mean, a trillion words encompasses virtually every Wikipedia article in existence. Every single one. Pretty much. Plus a massive chunk of the entire public Internet. Millions of published books, academic papers, scientific journals.
Legal contracts. Decades of court opinions. All of it. So engineers dump this trillion word ocean of text into a massive supercomputer. And they spend months.
Right. And hundreds of millions of dollars. Yep. In what they call the training phase. The training phase.
And the computer just churns through all that text. Exactly. Adjusting those billions of volume knobs. Learning exactly which words usually hang out together when humans discuss real estate or title insurance or literally anything else. But that training phase creates the foundation model.
So by the time you, the title professional, log into the software on a Tuesday morning, all of that heavy lifting is long gone. It's ancient history. You are interacting with the second phase. Which is called inference. Right.
Inference is the actual moment of magic. That's when you hit enter on your keyboard. Right. Right. During inference, the model takes your prompt, runs it through that massive mathematical map of parameters in just a fraction of a second, and generates a response for you.
But it doesn't just spit out a pre-written paragraph, does it? No. It generates the response token by token. Token by token. Yeah.
And a token is roughly equivalent to a word. Or maybe a sentence. It's ancient history. You are interacting with the second phase. Which is called inference.
Right. Inference is the actual moment of magic. That's when you hit enter on your keyboard. Right, right. During inference, the model takes your prompt, runs it through that massive mathematical map of parameters in just a fraction of a second, and generates a response for you.
But it doesn't just spit out a pre-written paragraph, does it? No. It generates the response token by token. Token by token. Yeah.
And a token is roughly equivalent to a word. Or maybe a syllable of a word. It is literally calculating the mathematically most probable next token. Over and over. Over and over.
At lightning speed. I love explaining it to people as smartphone autocomplete on steroids. That's the best analogy. Because when you type happy into your phone's text messages, your phone suggests birthday. And it does that because it has learned your personal texting patterns.
Right. Well, an LLM is doing the exact same mathematical prediction, except it is trained on a trillion words of human history instead of just your weekend group chat. And instead of predicting one single word, it's predicting entire complex paragraphs. So it's essentially a really well-read, mathematically gifted parrot. A very, very smart parrot, yes.
Because it doesn't actually know what a warranty deed is. It has no actual comprehension of property rights. None whatsoever. It just knows exactly what sequence of words usually follows the phrase warranty deed based on its training. Exactly.
And the prediction is so incredibly sophisticated that it simulates human comprehension perfectly. It really fools you. It does. When you read the output, your brain tells you that you are interacting with an intelligent entity that understands you. But behind the screen.
It is just pure high speed statistical probability. But that mechanism, you know, that statistical prediction brings us to the absolute most dangerous trap for anyone using this in a professional setting. A danger zone. Yeah. If the model is just predicting the mathematically most likely next word, what happens when the most likely sequence of words is actually factually wrong?
That is the core issue. Because an AI has no concept of truth or lies, right? It only understands statistical probability. You have hit on the single most critical limitation of this technology. When an AI generates information that sounds completely authoritative and 100 percent correct, but is entirely fictional, the industry calls it a hallucination.
A hallucination. Let's look at how that actually happens mechanically. Suppose you ask the AI to summarize a specific obscure county court ruling regarding a property boundary dispute. OK. The AI doesn't have a database to realize it doesn't actually know the answer.
It can't just say, I couldn't find that. Right. Instead, it just starts predicting. It knows the word Smith often goes with V, which goes with Jones. And it knows those words are usually followed by a volume number and a page number and a year.
Exactly. So it spits out Smith v. Jones, 452 F Part 3D 918. Wow. And it looks like a flawless legal citation.
Flawless. But it is a complete fabrication. It didn't look up a real case. It just strung together the characters that statistically sequenced together in its training data. So it will quote the wrong recording fee for a specific county.
Yes. Or invent a real estate statute out of thin air. All the time. And it will do it with absolute, unwavering confidence. That is terrifying.
And here is the brutal reality that every professional needs to internalize. Hallucination is not a bug. It's not a glitch. No. It is not a glitch that engineers are going to just patch in the next software update.
It is a fundamental core feature of how a prediction engine works. Because it's fundamentally creative. Yes. The exact same creative mechanism that allows the AI to draft a beautifully worded email from scratch is the exact same mechanism that causes it to invent a fake court case. You simply cannot separate the two.
Which begs the obvious question. I know what you're going to ask. If this thing just makes things up, has no concept of the truth, and lies with total confidence… Why on earth would a title agent or escrow officer trust it anywhere near legal or financial documents? What's fascinating here is, the solution is to completely reframe our expectations of the tool. Reframe how?
For a title professional, you have to mentally separate your entire daily workflow into two distinct buckets. Okay. Two buckets. You have your high-trust tasks and your verify-always tasks. You must treat this technology as a highly capable, incredibly fast first-draft machine.
A first-draft machine. Okay. That completely changes the expectation. Doesn't it? Because if it's a first-draft, you inherently know it needs human review.
Exactly. So what goes into that high-trust bucket? High-trust tasks are areas where the LLM is extraordinarily powerful. Because they don't require the AI to pull a hard, verifiable external fact. Okay.
So no specific dates or laws. Right. They just require the manipulation of language. Like drafting complex emails. Summarizing a 50-page thread of correspondence.
That's a huge time saver. Brainstorming marketing ideas for a new real estate farming campaign. Reorganizing messy meeting notes. Taking dense, jargon-filled text and rewriting it to sound approachable. The machine shines there because you aren't asking it for the truth, are you?
You are just asking it to organize information that's already there. Exactly. But then you have your verify-always tasks. The dangerous stuff. Like legal citations, exact recording fees, financial calculations, regulatory deadlines, contact information for specific escrow officers.
Right. If you ask an LLM to generate those hard facts, you are essentially playing Russian roulette with your closing. You really are. It is an incredible engine for momentum to get you past the blank page, but it is never the final word on facts. So once you establish that framework, the question then becomes, which specific machine should you use?
Right. Because there's a lot of them out there now. Because the landscape in 2026 is dominated by four major foundation models. And choosing the right one depends entirely on your specific workflow and, of course, your security constraints. Security is huge in this industry.
Let's start with the one everyone knows, OpenAI's ChatGPT. I mean, they pioneered the mainstream movement. They did. And it remains an absolute powerhouse for general-purpose tasks, brainstorming and writing. It's the standard for a reason.
But for professionals dealing with massive documents, you really want to look at Anthropic's model, which is called CLAUD. Yes. CLAUD is renowned for its focus on safety and careful reasoning. But its real superpower for a title professional is its context window. Let's explain what a context window is because that is a massive differentiator.
Think of a context window as the AI's short-term memory. Short-term memory. OK. Older models had very small context windows. If you fed them a 50-page document, by the time the AI reached page 50, it had completely forgotten what was on page 1.
That's not helpful for a title search. Not at all. But CLAUD has a massive context window. It can hold a 200-page commercial contract in its head all at once. Oh, wow.
It can instantly cross-reference a clause on page 198 with the definition buried way back on page 3. For an escrow agent reviewing massive files, that is invaluable. It's game-changing. And you have Google's Gemini, which is brilliant if your agency is already running heavily on Google Workspace. Because it's deeply integrated right into Gmail and Google Docs.
Super convenient. But I want to spend some real time on the fourth major model because it solves perhaps the biggest headache for our industry. Meta's model. Yes, Meta's model called LLAMA because LLAMA is an open-source model. And open-source completely changes the security paradigm.
It really does. In the title and escrow industry, you are handling highly sensitive, non-public personal information. NPI. Social security numbers, wire instructions, financial histories. You absolutely cannot take a document containing a client's wire instructions and paste it into a public web browser using OpenAI or Google.
Never. You are sending sensitive data to a third-party server, which is a massive compliance violation. But because LLAMA is open-source, the actual underlying code of the mathematical model is freely available. So your agency's IT department can actually download LLAMA? Yes.
And install it directly onto your own internal lockdown servers. So the data never leaves your building? Never. You get all the predicted power of a massive AI, but you retain complete control over the privacy. For a heavily regulated industry, locally hosted open-source models are kind of the Holy Grail.
They absolutely are. Now that we know the models and we know the security constraints, let's look at what this actually looks like in practice. Our source material outlines some incredible real-world workflows that title professionals are using right now in 2026. Real use cases. Yeah.
Let's say you just finished a brutally complex commercial closing in Chicago. It took 3 hours, emotions were high, and there was a massive dispute over a zoning exception. Now you have to write a follow-up email to the opposing counsel summarizing the resolution. I'm exhausted just thinking about it. Right.
Normally, the mental friction of just getting started on that email, you know, making sure the tone is professional but firm, hitting every detail, that takes 15 minutes of staring at a blinking cursor. But using the LLM as a first draft machine, the professional just types a messy stream of consciousness into the prompt. Just word vomit. Literally. Draft an email to Attorney Smith.
Chicago commercial closing is done. Mention we agreed to hold $50,000 in escrow for the zoning exception. Make the tone highly professional but keep it brief. And in 3 seconds? In 3 seconds, the AI generates a beautifully structured, perfectly polite email.
The professional just reviews it, maybe changes one adjective, and hits send. A 15-minute email. Using the LLM as a first draft machine, the professional just types a messy stream of consciousness into the prompt. Just word vomit. Literally.
Just draft an email to Attorney Smith. Chicago Commercial Closing is done. Mention we agreed to hold $50,000 in escrow for the zoning exception. Make the tone highly professional, but keep it brief. And in three seconds?
In three seconds, the AI generates a beautifully structured, perfectly polite email. So the professional just reviews it, maybe changes one adjective, and hits send. A 15-minute chore becomes a 45-second administrative blip. Or look at title search summarization. An examiner receives a bloated, 200-page title search full of boilerplate language.
Standard Tuesday. Instead of manually reading line by line to find the single mechanics lin, they feed the document into a model with a large context window. Like Claude, like you mentioned earlier. They prompt it, act as a meticulous title examiner. View this document and extract only the breaks in the chain of title, active lins, or unusual encumbrances.
Present your findings in a bulleted list. And within seconds, the AI filters out 198 pages of noise and highlights the two pages that actually matter. It is also transforming prospect research for business development. Oh, this is a great one. Say a sales rep has a lunch meeting in an hour with a high-producing realtor they've never met.
They ask the AI to synthesize all publicly available data on this realtor's recent sales, the demographics of their brokerage, and the current market trends in their specific zip code. And they get an instant, highly-customized briefing document to just read in the Uber on the way to the restaurant. Incredible. And my absolute favorite application, client communication. Crucial skill.
Taking a dense, archaic title exception written in just awful legalese and asking the AI, explain this title exception using simple analogies so a nervous, first-time homebuyer can understand exactly what it means for their backyard. That is so powerful. And here's where it gets really interesting. Not a single one of those examples requires you to know how to code. The barrier to entry has completely vanished.
The only skill required is the ability to speak plain, clear English. Which is a skill every successful real estate professional already possesses. Right. But the single most important technical skill in 2026 isn't programming, it is mastering the art of prompting. Prompting.
Prompting is simply the way you communicate your request to the AI. If you give a vague, lazy instruction, the statistical prediction will generate a vague, lazy output. Garbage in, garbage out. Exactly. But if you provide context, specify a persona, and clearly articulate the desired format, the model will do the heavy lifting beautifully.
So arguing over whether Claude or ChatGPT is technically superior is kind of missing the point. Entirely. The value comes from how effectively you can steer the machine. Okay. But if the barrier to entry is really just speaking plain English, why is it so urgent for the people listening right now to start experimenting immediately?
Good question. I mean, why not just wait a year or two until this technology is automatically baked into the title production software they already use? Because the title and settlement industry has historically been very slow to adopt new operational technology. That's putting it mildly. Right.
And that institutional sluggishness means the window of advantage is wide open right now. Right now in 2026. The professionals who are leaning into the discomfort of learning this today, while their competitors are still skeptical and waiting on the sidelines, they are building a compounding competitive advantage. And a compounding advantage means every single week, you use it, you refine your prompt, you figure out new workflows, and you get a little faster. Meanwhile, the person who refuses to use it is stuck at the exact same baseline speed they were at in like 2015.
Precisely. AI is perfectly designed to take over the time-consuming mechanical production of knowledge work. The synthesizing data, drafting boilerplate text, organizing research. But it absolutely does not replace the title professional. No.
No, it liberates them. If the machine is handling the mechanical drafting, the human being can focus 100% of their energy on the things that actually differentiate them in the marketplace. Like their specialized judgment. Yes. Their localized industry relationships and their deep expertise in navigating the emotional friction of a difficult closing.
An AI can't read a room. Exactly. So what does this all mean? I think it means we have to acknowledge a very harsh truth about the future of this industry. Let's hear it.
Professionals at risk of losing their jobs over the next five years are not the ones who are going to be replaced by an artificial intelligence. No, they're not. They are the ones who are going to be radically outpaced by their own colleagues who have been quietly using AI as a first draft machine to do the work of three people. This raises an important question. How long can you afford to manually draft emails and manually read 200-page boilerplate documents when the boutique agency across the street is doing those exact same tasks in 14 seconds?
You just can't compete with that. You can't. Adapting to this isn't just about learning where the buttons are on a new piece of software. It is about recognizing a fundamental tectonic shift in how value and speed are generated in your profession. Wow.
We have covered a massive amount of ground today. We really have. We completely dismantled the idea of software as a rigid vending machine. We learned that a large language model is actually a pattern-based prediction engine, you know, an autocomplete on steroids trained on a trillion words. Not a literal librarian.
Right. We've seen how it functions as the ultimate first draft machine that can instantly summarize a chain of title or draft a complex email, saving hours of tedious mechanical work. Just massive time savings. And most importantly, we learned why it hallucinates. And why you must ruthlessly separate your daily tasks into those high-trust and verify-always buckets.
Yes. How do you truly unlock all this potential? The immediate next step is mastering that communication layer. The prompting. Right.
The art of prompting is what separates a mediocre AI parlor trick from a workflow that genuinely saves you 10 hours a week. It all comes down to the clarity, context, and structure of how you ask the question. It is entirely about how you steer the machine, which leaves you with a final provocative thought to chew on as you go about your week. Okay. So we were talking about how comfortingly predictable technology used to be.
But now we are talking to mathematically gifted prediction engines that write original thoughts based on statistical patterns. It's a whole new world. It really is. So if artificial intelligence completely takes over the mechanical production of knowledge work, the drafting, the summarizing, the data synthesis, how will the fundamental definition of expertise change in fields like real estate, escrow, and title insurance? Over the next decade, will your professional value shift entirely away from knowing all the rules and move entirely toward asking the machine exactly the right questions?
In a world where change is the only constant, Mo Shumil stands at the forefront, guiding title professionals to not just grow their businesses, but to master the art of innovation. With every episode, you're handed the keys to unlock unparalleled growth and stay ahead of the curve. Get ready for a transformative journey. Mo Shumil www.mo-shumil.com
