Episode 11 ·
Agentic document extraction, visual AI & enterprise trust
Dan Maloney — CEO at Landing AI
Dan Maloney, CEO at Landing AI, talks about agentic document extraction, visual AI and trust on the Future of Data and AI Podcast, hosted by Raja Iqbal.
Transcript
Chapter 1
Introduction
Raja introduces Dan Maloney, CEO at Landing AI, and kicks off the episode.
Raja Iqbal: My guest today is Dan Maloney. Dan has spent over two decades at the intersection of enterprise software and AI, founding companies, selling them to the likes of Cisco and DataRobot, and building go-to-market engines at SAP and Sapphire Ventures. In 2024, Andrew Ng decided to step back from the day-to-day and Dan was appointed the CEO of Landing AI. Landing AI is on a mission to solve one of the most stubborn problems in enterprise AI: ingesting and understanding the world’s documents and visual data. Welcome to the show, Dan.
Dan Maloney: Thank you, Raja. I appreciate you having me here.
Chapter 2
“LASIK for LLMs”
Dan unpacks the analogy he uses for Landing AI — why LLMs are good but fuzzy at reading documents, and why at enterprise scale that fuzziness is extremely costly.
Raja Iqbal: So Dan, let’s get started with one of your comments that I’ve seen in one of your interviews that I was going through. You have referred to Landing AI as a LASIK for LLMs. Tell us what you mean by this.
Dan Maloney: First of all, again, thank you so much for having me on the show. I appreciate it. I’ve watched in the past and I think it’s a great show.
When I talk about the LASIK for LLMs, it’s really just an analogy I try to use for a lot of people who are familiar with working with AI, working with LLMs, but are really used to using it more for tackling unstructured data much more in the world of text and the language dynamics only.
I really think about Landing AI bringing a lot, because we have deep expertise in visual AI. We do a great job at taking specifically images, documents, other elements that are more visual in representation, and really bringing clarity and accuracy to them.
It’s not that you can’t give an LLM a document to process. It’ll do an okay job at that. Sometimes it’ll be 70, 80, 90 percent accurate. Other times it won’t be quite as high. But when I think of LLMs, they do a great job at guessing what that next token is. For example, if you ask an LLM, “Hey, what did you do this morning?” If you ask a human that, and then I ask you the same question later in the day, it’s generally going to be about the same, but it’s not going to be 100 percent accurate. The LLM is the same way. You could almost ask it the same question and it’ll give you mostly the same thing.
But in our world at Landing AI, our broader vision is making the world’s documents computable. If we’re dealing with hundreds of millions to billions of documents for some of our largest customers, just that little bit of error rate, just that little fuzziness, if you don’t have that LASIK for LLMs, if you’re not 98, 99, 100 percent accurate, it’s extremely costly to them. Many times too many humans will have to be involved in a process they don’t want to be involved with.
So I feel like if we can crack the code and be the LASIK for LLMs and really give them that crystal clear accuracy in taking that unstructured data and making it structured in a 100 percent accurate way, or aiming for that, that’ll really make a big difference in the world. That’s where the LASIK for LLMs comes from.
Chapter 3
Why OCR Struggled for Decades
Dan explains what OCR always handled well and where it fell down, and why Landing AI chose to come at documents from computer vision and agentic reasoning rather than building on top of OCR.
Raja Iqbal: OCR has been around for decades now. So what has happened in the last couple of years? What has enabled AI to be able to crack messy, complex documents like tables, and tables within tables, and gridlines, and invoices scanned at odd angles, handwritten invoices, and all of that? What has happened?
Dan Maloney: This was a problem I was dealing with back in my days at SAP, dealing with some of the biggest enterprises, or really just all enterprises. It didn’t matter what the company was. Every company has their accounting and finance department dealing with invoices and receipts. You have more verticalised use cases like know your customer in financial services, revenue cycle management, or prior authorisation, these things in healthcare and life sciences.
Where OCR did a great job traditionally was when you have a simple document, mostly text-based, reading left to right. Sometimes handling multiple languages. When you had multiple different languages that were quite diverse, if you had English and maybe some European-based language, an Asian-based language, generically speaking, they might struggle. Some got better over time by putting those pieces in.
But really, OCR did a great job at the simple PDFs, digital native. They really struggled with the messy PDFs, multimodal. You might have text chunks, complex tables, figures. What do you even do with a figure? Multiple different languages.
Raja Iqbal: And images. And graphs with annotations.
Dan Maloney: That’s right. Annotations. You’ve got scanned documents. I think of the US government releasing the JFK files. Those documents are from the 1960s. How well were they kept? They were a scanned copy. So all of that complexity, OCR mostly stayed away from. They handled the simple use cases.
When we started looking at this problem, I remember Andrew and I a few years ago were looking at it with our AI researchers who have deep understanding of computer vision and visual AI more broadly. We said, I really believe with some of the approaches we’re starting to see around agentic reasoning, and starting to look at what the LLMs were doing, VLMs, and some of the custom models that we were building, we really think we can come into this problem not from OCR, not growing up from OCR and making a few modifications.
OCR did an okay job, but it was very schema-driven, building a lot of templates to handle the different cases. There’s a lot of manual work, heuristics, coding, to handle complexity. We really felt like coming at it from a generative AI or machine learning point of view, with some of the newer techniques around agentic reasoning and agentic approaches, could help us solve problems that weren’t solved.
We had been building a lot around a stack to tackle images, documents, other elements. We put out a very rapid prototype internally and we said, let’s go for it, let’s see what we can do. We threw some of the most complex documents at it. This was probably a couple of years ago. It did an exceptional job. It was also really costly behind the scenes for us. We were like, this is great, we actually solved the problem. The cost is too high, but we’ve solved the problem.
So we actually put that out to customers. We didn’t charge the customers early on for all the costs that we had. We said, we know that we’ve built something great, we can solve this problem, and over time we’ll drive the cost down.
That’s a lot of what we’ve been driving new innovations around, looking at cost efficiency for customers. One of the main things that’s slightly different: when you think about the frontier labs, what OpenAI or Anthropic or Google and others are tackling, they’re trying to tackle the problem of intelligence. AGI. That’s why you see new models coming out, and new models are more expensive. The price keeps going up.
We have to deal with a slightly different world. In tackling the world’s documents and making them computable, we’re focused on high accuracy, dealing with the complexity of the different modalities in a document, whether digital or scanned, all the different complexities within a document: the tables, the text chunks, the figures, the signatures, the attestations, the stamps, all of the things you know about.
But we need to continuously look to drive costs down, to be able to make this a problem where we can basically take the human out of the equation. When we talk to so many customers, and the analysts that are involved with this, they don’t want to be involved. Some of the big banks we talk to have thousands to tens of thousands of people involved in doing attestation, checking, making sure the data going over is right, because if they make a mistake it can be worth millions of dollars, those individual mistakes.
So we’re trying to crack the code on this really unique approach to solving this problem. But we can’t just keep driving the cost up, because when you multiply that across billions of pages, every cent that we can save per page is massive savings to the end customer.
It’s a really interesting approach, how we take the blend of building out some of our own models that are purpose-built. For example, we have our Document Pre-Trained Transformer. We’ll go into that a little bit later. We have some of our own proprietary models that we’re using for transcription and other aspects, and we blend in the use of LLMs when needed. This blending of intelligent routing, which models to use, I put all of this into our document harness around our inference. Cracking the code on all that is what really makes this problem unique. It makes it challenging, but we believe the tech I’ve just talked about makes it so that we’re finally getting close to cracking the code on something where, as you said, OCR has been 10, 20, 30 years at different companies trying to solve this problem.
Chapter 4
Faster, Better, Cheaper: Multimodal Models vs Landing AI
Dan breaks down the three vectors customers actually evaluate, then explains why models are a surprisingly small part of the equation — and introduces the DPTx family and Landing AI’s document harness.
Raja Iqbal: We also have multimodal models. Let’s say as a consumer, or as a potential customer, I’m trying to make a decision between using a multimodal model or going for Landing AI. Is it cost? Is it accuracy? Is it speed? How would you differentiate?
Dan Maloney: Even maybe taking Landing AI out of the equation for a moment, when companies look at this problem, they want something faster, better, cheaper. By the way, that’s pretty much true across all products when you’re purchasing.
So speed absolutely matters in certain cases. When you’re dealing in the realm of document processing, there are definitely use cases where a consumer or your customer is involved. Let’s say you’re dealing with your bank and you upload a check and you want it to scan the check and verify your signature or pull an amount off it. I’m not willing to wait minutes to hours for that response. So in those cases that are more consumer or customer-facing, where speed matters, that will be critical.
That’s not all document use cases. Sometimes we go out to customers and they’re like, hey, we’ve got a billion documents that have just been sitting there for the past 100 years. We want to get that oil, or gold, or that data out of it. How do we take that dark data and turn it into something valuable for us? And it’s okay if it runs in batch mode and we get it in hours to a day. That’s fine. So speed does matter, but it depends on the use case.
Then, better. I think better is the accuracy. There’s no doubt that every percentage point of accuracy higher that you can get for a customer absolutely has an impact. When we first brought our product to market, it was super accurate. It wasn’t necessarily the fastest in the market, and it wasn’t the cheapest in the market, but it was solving use cases that had never been solved before, so that was really valuable.
The last piece is that cost equation, and it really depends. Some use cases you’re willing to pay more for, per page or whatever your metric is. For things like receipts or invoices, things that are very common and high volume in a company, that aren’t necessarily so high value, you want to do that at the best price possible, as long as the accuracy is there.
When companies look at those vectors, depending on the use case, they look across the three and say, okay, I’m going to make a decision on either one platform or multiple platforms, one model or multiple models.
One thing I’ll say now, taking it back into the Landing AI world: models are an important part of the equation, but they’re actually quite a small part of the equation in the bigger picture. First of all, we use a family of different models to solve different problems. We’ve built some proprietary models, what we call our DPTx family, our Document Pre-Trained Transformer. We have different generations. We’re actually coming out right now, as we speak, with our DPT-3 generation.
What these purpose-built models do, for example our DPT model, is a great job at taking a look at a document almost like a human does, or a page almost like a human. It quickly looks at it and understands the layout. It says, I see a text block, here’s a table, here’s an attestation, a signature or something along those lines, some marginalia like page numbers. It very quickly does that, and it breaks things down at a block level, might even take a look at a line level, down to a word level.
Then we have other models that take it downstream from that. We’ve got a model that’s purpose-built to handle text very well, high accuracy, can do it super fast and super cheap. Versus maybe we want to use a different model for something more complex like a figure, and being able to explain that. There might be text on that figure, but there also needs to be a generation of some text to describe what you see.
So there’s all these elements that we put together. I like to think of it as: we’ve got the models, which play an important part at the inference level. We have our harness, and when I think about our document harness, it handles things like the Document Pre-Trained Transformer, our intelligent router, preconditioning parts of the document for an LLM or other models downstream. And then on top of it we have our tools or APIs that do things like taking a document and parsing it. A parse is when you literally take a look at a document and turn every part of it into structure. Then you might extract key-value pairs. For example, I want to grab a name out of that. I want to grab Raja out of that document and put it into my downstream system, into my SAP system, into my Snowflake data warehouse or Databricks, whatever it might be. Or I want to take that document and throw it into a RAG system downstream to be able to ask questions of it.
So there are so many downstream dynamics. When you start looking at what we provide, the models are an important piece of the equation, but we’re providing an overall agentic system that’s purpose-built to handle that holistic intelligent document processing, which again can leverage the latest and greatest models. If a new open source model comes out, or one from the frontier labs that are more commercial, those can be plugged in. But those are only a very small piece of the equation, and we’ll only send to them when it’s relevant, when it’s necessary. Normally they’re more costly, so they’re normally a very small part of the overall answer back to a customer problem at the end of the day.
Chapter 5
Plugging Landing AI Into Your Own Stack
Raja asks whether you can simply drop Landing AI into an existing pipeline for ingestion. Dan explains why the product moved beyond complex documents to handle simple ones too, and introduces complexity-aware pricing.
Raja Iqbal: So basically, if I have a scaffold, I have a harness that I’m using, I can simply bring in Landing AI for my document ingestion, and then just do whatever my RAG is. My MCP, RAG, inference, whatever I’ve been doing, but bring in Landing AI for complex documents.
Dan Maloney: That’s right, and really for all documents. We can handle the complex documents, but a lot of companies over the years have been dealing with a lot of different tech.
One of the biggest requests from a lot of our customers was, hey, we really like Landing AI, we like your Agentic Document Extraction product, but it’s only handling the complex documents. And that’s what we launched with. That was our wedge into the market. But a lot of them were saying, we love the product, we love the accuracy, we love the visual grounding that it provides, it handles the complexity of tables, breaks things down to the field level, all this is great, but we have to go use another product when we want to handle invoices or receipts or things that are digital native.
So recently we’re bringing to market the ability to put those all together. Whether you’re handling something complex or simple, we can handle it in our platform, and we can also do that cost adjustment on the fly. Where you need the complex dynamics, we can charge you more. We have something we call complexity-aware pricing. When we understand that it’s more complex, we’ll charge appropriately, and when we can do it for almost OCR prices, we can do that for your simpler documents.
At the end of the day, when document processing is part of your finished product or your end application, you can call our APIs and let that become the engine underneath that powers it. Whether you’re building a product to deliver to customers, or using it internally in an enterprise. A lot of the biggest banks in the world in financial services, some of the biggest healthcare companies, legal, and many more verticals use this. But we’ve had a big pull-in from financial services, healthcare, life sciences, and legal, because those industries are so document intensive.
Chapter 6
Measuring ROI on AI Initiatives
Raja asks how companies should quantify the value of AI spend. Dan argues document processing is unusually measurable, then steps back to explain what the token maxing era was actually about and why CFOs started tightening the belt.
Raja Iqbal: Thanks, that was very helpful. There’s a question that has bugged me lately. I’ve been involved in building products, not just on the ingestion side, but ingestion is a part of it. An end-to-end platform that we have built. How do you measure ROI on AI initiatives? How do I quantify this whole concept of token maximisation, token minimisation, tokenomics? How do we measure this stuff? I know you have been in industry for a while, you’ve been an entrepreneur. At the end of the day it is about adding value to the business. How do we quantify?
Dan Maloney: Listen, I have my thoughts. One of the benefits of my world specifically, so maybe I’ll talk about my world for a second around Landing AI and document processing, and then I’m going to take a bigger step out.
One of the beautiful things about what we’re doing is that it’s very quantifiable. Our work is very outcome-based. A customer has a document that they want to process. One document, 10 million documents, billions of documents. They know what they want to extract. Whether you’re putting something into a RAG, having an agent do some work on it, taking key-value pairs and driving that into mortgage underwriting, whatever your downstream use cases might be.
If we can do that at a high accuracy for them, they know that when we process a document, that is a business outcome for them. And they can also see the accuracy of what we do. So they can say, today we’ve been doing this with thousands of humans involved in this process, because we have to have them. Even though these are analysts who do other things, we also have to do this.
Even though these humans do not want to be doing this, looking at millions of documents, if you’re only 80, 90, 95 percent accurate, it’s almost like being 0 percent accurate, because they still have to go back and validate so much. And a lot of times you can’t tell them exactly where in the document that exact field is that you’re not sure about, so they have to go back and read a bunch more than they normally would.
So our ability to do that with that high level of accuracy, and we’ve actually processed that document, that is an AI outcome where customers say, great, we’ve processed these documents, these number of pages, with this high level of accuracy, we have low error. That’s really valuable, and that is a measurable dynamic where they can tie it to either removing costs out of the equation, or even sometimes new revenue models, because now they can process online and do things that used to take offline analysts to do and come back to a customer with an outcome. Now you can sometimes do that in real time. Approve a customer for a credit report, approve that mortgage, whatever that might be. So it can have a big cost saving or revenue outcome.
In the world of documents, applying AI to that is a little bit more obtainable. And you can do a lot of benchmarking and evals around it to see the improvements you’re getting. A lot of companies, because they’ve been doing this for 10, 20, 30 years, have a lot of in-house metrics and benchmarks. That’s what’s exciting about this and why customers are paying for it and see value in it.
When you take a step back from that, it makes sense to me where maybe in 2025, maybe end of 2024, 2025, and even a little bit in 2026, you were seeing things around token maxing. It’s not so much that it’s about maximising your tokens for a moment, but you’re trying to see, is there a way we can do things so radically different that maybe we can bring product to market faster, we can out-innovate our competition?
Just like I didn’t want cost to be a limitation up front for my team. I said, see if you can solve that problem, then we’ll work through cost cutting. Last year a lot of companies were doing that. There wasn’t an unlimited AI budget, but they were almost rewarding you for taking a look at using Claude Code, Anthropic, OpenAI, Google, whatever it might be, or even internal tools, to see if that unlocked things that were a little more ambiguous and hard to define. They were like, let’s see if the innovations come out. And if something revolutionary came out, great.
But the CFO is also saying, after a year of maybe not seeing anything breakthrough and tangible, and you’ve spent millions to tens of millions to hundreds of millions, now we’ve got to tighten the belt a little bit. Any time new innovations come out, that’s not an uncommon dynamic, where there’s an overreaction one way and then you start bringing it back.
So what you’ll start seeing now, out of that work done over the past 12 to 18 months, is what were the signals in the noise that might have shown promise. I think documents is a place where a lot of companies saw benefit.
And I’ll tell you this: companies are growing. Their revenue is growing for a lot of companies. Their employee counts, though, aren’t necessarily growing. They’re saying, we’re shifting some budgets more into AI, we might have to do more with less. They’re doing more business. For example, with a lot of banks, they’re saying, we’re growing our customer base by 50 percent year over year over the next three years. Everything that goes along with that, the document processing, is going to grow at a similar level or higher, but we’ve got to do it with the same amount of people or less. That’s where AI and automation come into the equation.
So you’re seeing companies really look at different use cases where they can’t scale with people anymore, because they’re at their max limit of what the economics will let them do. That’s where they’re trying to apply AI, and then build evals around that to see how well it’s performing.
Chapter 7
Why Document ROI Is Clearer Than General AI ROI
Raja draws out the contrast between measurable document pipelines and fuzzy general AI usage. Dan introduces the framing he uses for the product: autonomous, auditable document processing.
Raja Iqbal: As you were explaining this, it suddenly became clear to me that with automated document ingestion and processing, the unit economics are established and quite clear. Previously we had 10,000 claims a month coming in and a staff of 100 people, and now we can process the same number of claims with 20 people. And I know the outcome, whether the claims are processed correctly or not. So I have a pipeline where I inherently understand how valuable that AI usage was. Versus if everyone is now writing their documents, their product plans, their emails, or for that matter code, using AI, the measurement of ROI is going to vary from case to case.
Dan Maloney: That’s right. That’s one of the benefits of the world of Landing AI and applying AI to it. There’s a lot of measurement that’s been done.
As I often talk about what the product does itself, I often talk about autonomous, auditable document processing. The autonomous part is allowing us to help those companies not have to add more and more humans to solve something that’s very tedious. Their analysts, or whoever’s solving that challenge, can now focus on the solution they’re actually providing. Their end solution is not document processing. They’re delivering some solution at the end. I don’t care whether it’s mortgage underwriting, invoicing, contracts, legal, whatever it might be. They can focus on that and they don’t have to worry about the accuracy piece.
And all they want to be able to do, if it’s fully autonomous or mostly autonomous, is have an auditable check. If I gave you a billion documents, I want to be able to go back to that document, to that page, to that chunk, to that field that you said is there. If I’m not sure, I want to validate that really quick. As long as you have those two things working hand in hand, it’s a beautiful thing.
But you’re absolutely right. This has been a thing not just with AI. I remember when the internet came, and then we went in big with SaaS, and we started building apps, and it was eyeballs back in the day in 2000. Just having eyeballs to your site. Everybody’s always playing with different metrics, and defining that North Star is a very difficult thing, especially when you go down to something like an AI OS. When you’re saying, oh, you can use it for email and writing your emails, and by the way it can build an agenda for me today, and then you’re like, oh, time savings, I get that back. Those things you’re going to struggle with.
So I always say to the team, exactly as you were calling out, it is important to measure. It’s going to be very important to measure when you want to go to the CFO and ask for budgets. If you can be very good about what AI is doing for you, and you know what it doesn’t do well, that’s the case where you can then say, all right, we’ve maximised what AI can do for us here, and by the way, we still need additional marketing budget for this, or we might need some human resources to come into the equation.
There are so many things that AI is tackling, from complex physics problems and quantum mechanics down to, like you said, helping optimise your email inbox. And everything in between. So everybody, depending on what you’re tackling, has to think about what your North Star is, what your key metrics are, what you’re selling to your customers, and making sure it’s impacting their business. It’s not one metric, one thing fits all.
Chapter 8
Token Maxing and the Industry’s Correction
Raja raises the pullback at Meta and Microsoft after a period of unrestrained AI spend. Dan argues loosened guardrails are a normal response to new technology, and explains why Landing AI deliberately isn’t token maxing.
Raja Iqbal: Do you think the industry is coming toward that point, or at least starting to converge toward it? Because initially there was this mad rush. Hey, everyone, just go and get all the tokens that you can use. I think two days ago there was an article about Facebook’s token usage, their token budget, that estimated something like $5 billion per year if they just let it continue.
Even companies with a very mature understanding of these technologies, and no one should be able to understand that better than Meta and Microsoft, even they were first to let it go, unleash everything, go crazy. And suddenly now they’re rolling back or curbing the usage of these tools. So do you think the financial leadership and the technical leadership is now realising that maybe initially they miscalculated?
Dan Maloney: When new tech comes out, I don’t know if it’s a miscalculation. When you don’t understand what the tech can do, maybe sometimes you loosen the guardrails to start. You want to see what comes out.
Now, on the flip side, if you’re spending more than you expected: okay, finance didn’t quite understand, maybe even the tech didn’t understand. Oh my gosh, now I can run agents, I can spin up 100 agents, 1,000 agents. I can put them in a loop to solve problems.
We often talk about humans and 996. If you’re working really hard, you’re working nine to nine, 12 hours a day, six days a week. Agents can work 24/7, 365. My son told me that MacBooks, after something like 48 days of running continuously, need to be shut down. So there are new things that we’re discovering. If you have your agents running locally, you have to reboot.
But you’re spending more, and potentially a lot more work can be done at those times and get you ahead of the equation. But like you called out, if you’re not solving something, and it’s not driving toward that solution or getting ahead of the competition, then it’s a big negative.
On the flip side, if you spend $5 billion in a year and you change the game, and that’s going to drive $100 billion for you over the coming year, no CFO is going to complain. But they’re going to complain when they’re not seeing the outcomes.
And I’ll say this, this is not an AI challenge. This has been a challenge back when things went from on-prem to SaaS. All of a sudden with SaaS you had all these apps. IT, your CIO, used to lock everything down. Then all of a sudden business units are buying things, and you have this shadow IT budget going out there that you didn’t even know about. AI is a better, faster version of driving awesomeness in that area. Now you can build your own apps, so you can do your vibe coding with your sales reps or your operations people.
We’re still in the early innings here, and I think everybody’s trying to find out together the best way to do this. On our side, when you’re using always the latest and greatest models, the cost for driving toward intelligence, like the frontier labs, gets higher and higher. We’re already very much thinking about how we solve problems, and then as soon as we solve it, how do we do it massively cost efficiently so we can pass that cost efficiency on to our end customers and still make the margin for us.
I think we’ve found our lane. We understand it. We’re not trying to token max in our area. We understand the problem quite well. We’re trying to do things massively efficiently, use smaller models where applicable. It’s not quite a mixture of experts, but we use an intelligent router to understand the different documents, pages, blocks within it, and distribute those accordingly. That’s what we’ve done and how we’ve tackled it in our world.
Chapter 9
When Broken Processes Limit What AI Can Do
Raja argues AI can only go as far as an organisation’s processes allow. Dan agrees, then describes the shift from human-first to agent-first design and why even Landing AI had to re-engineer its own processes.
Raja Iqbal: Do you think, Dan, in many ways, and I like this analogy: if the roads don’t have lanes marked, autonomous vehicles can only work so well. You have to have the lanes. Very often when I look at customers, and when we look at problems internally as well, one of the challenges is our operating model, our processes. If our processes are broken, our operating model is broken, and in the technical lingo, if our context is broken, then AI can only do so much.
Do you think this correction mechanism will eventually evolve into companies realising that if you want to be an AI-first company, you’d better go and fix your processes? Because when humans are involved, a human can step in if something goes wrong. But AI will either take a bad decision or it will fail. There are only so many options. Do you think that eventually, when companies start realising what’s wrong with their processes and their operating model, that is going to be fixed?
Dan Maloney: I think there’s a lot of merit to what you said. Until recently, all companies were built human first. Obviously we had humans, the human processes. How do we build things for the humans? And then the humans looked at automating. But even when we automated, it was still all human-driven.
I think it was really interesting, the other day I had a conversation where they were literally saying how it’s now kind of agent first. It’s like an agent-centred design. It’s agent first from the get-go, and the human comes in at key points for a lot of the tasks. They’re doing so much that in the old days might have taken thousands of people to do, that they can now do with 10 or 20. The agent-first approach is to drive that dynamic, not with the human first in mind, but the agent. So let the agent take a lot of the remedial tasks, and those go up, but then the humans can sit on top more with the strategies, the constructs. Yes, come in with the guardrails here and there.
Even a company like ourselves, which is a young company in the bigger picture, I’ll call us a late-stage Series A company, we’ve even had to evolve our process. Even though 90 percent of our company are strong AI people, AI has gotten so strong just in the past couple of years that even we’ve had to evolve into a stronger, better AI-native company.
So imagine if you’re a company 10, 20, 30 years old, with all your processes set up and everything built in that construct. Of course, as AI allows you to do things differently, that’s got to go back through your company. And that could have a ramification of reallocation of people, reshaping of the company. Some people and their roles might not be there. New roles might pop up. That’s what the companies that have been around a little bit longer are all taking a look at, just like we did with digital transformation, and it’s maybe even accelerated.
When I was back at SAP in the 2000s and 2010s, we were definitely talking about digital transformation for companies. But when COVID happened, everybody was like, oh man, we need to be fully digitised, work remote, et cetera. And now I think there’s another evolution happening here with AI and how that impacts things.
Andrew does a great job, of course here with Landing AI, but also doing education with things like DeepLearning.AI and other companies. When we go into companies, half the time we’re talking Landing AI and document processing. But sometimes we just step back with customers and talk about the impact of AI more broadly, using the use case of document processing as one of those. We often talk to them more broadly about the impact and what it’s done with us.
My engineering team was writing all their code three, four years ago. Right now, basically 100 percent of our code is written and reviewed with AI. And we run into a friction point where AI is not doing that, and then we say, okay, because this part is moving so fast, we hit this friction point. Then we look at how AI can solve that, and we just keep moving.
And that goes into not only engineering, but into product. If engineering is moving so fast, that puts pressure on product. Then the product management team solves that. That puts more pressure on the go-to-market team. So it’s this continuous cycle of the whole company speeding up. And we keep getting better.
That’s where, again, a lot of companies were token maxing, or benchmarking even, to show how well they’re doing. But they were doing it to transform the company, to hopefully get a competitive edge over other players in the market. It’s not an uncommon tack that was taken. With any revolutionary technology, you don’t want to be the one left behind. But you start getting controls around it, and if it gets too crazy on the spend, a company can’t survive if it’s not making more cash. If it’s burning more cash than it’s bringing in, ultimately that’s a problem. You can’t keep going to investors and VCs forever. It feels that way sometimes, but it’s not the case for most companies.
Chapter 10
From the Factory Floor to Documents
Raja asks about Landing AI’s origins in visual defect detection. Dan explains why the same spatial awareness and localisation techniques transferred directly to documents, and why documents were the bigger market pain.
Raja Iqbal: So Landing AI started on the factory floor. As I understand, Landing AI’s earlier products were more about detecting defects in the production line, and moved into agentic document extraction later. Do you see any connection between the two problems?
Dan Maloney: I would say Landing AI’s broader path is still the same. They started in trying to solve vision problems using visual AI and computer vision techniques.
One of the early things Andrew and the team were looking at, even before I joined the company, was manufacturing and industrial automation. You’re looking at something that’s getting manufactured, and you want to look inside a specific area and maybe look for a visual defect, a scratch, a chunk left out, whatever that might be. Or you want to identify, oh, there’s a screw right there, or a screw is missing.
So the spatial awareness, that ability to localise within an image, as well as understand and classify the image itself, those techniques are absolutely applicable across all visual AI. Whether you’re talking about computer vision for a car, looking at a road and understanding the lines on the road that help it stay within the lines. There are newer approaches where you don’t even necessarily need lines always there.
These same visual AI techniques that we learned there, we started applying to documents. You can give us a document literally as a digital PDF. You can also take a picture of a receipt and give it to us, no different than taking a picture of a semiconductor chip or whatever it might be.
One of the really nice things about the world of documents is that there are trillions and trillions of documents that exist out in the marketplace. You do not necessarily have to get the right lighting, lenses, and camera hardware in place to capture them. The documents, the formats, are already there. That’s a very powerful dynamic. It’s a very big pain that a lot of enterprises are dealing with, but they have the data. The only problem is it’s an unstructured world, and they’re just looking for something to help convert it so they can use it downstream.
So those same techniques that we learned around classification, around localisation, things like object detection, allow us to do things like visual grounding and that spatial awareness within a document, within an image that we receive. I think that’s really why we’ve disrupted the market, because we had that deep understanding.
And I wouldn’t say we focused solely in that arena of manufacturing and industrial automation. When Andrew was back at Google with Google Brain, and then over at Baidu and these other areas, they were looking for ways to solve real-world problems, visual problems, and not just tackle the world of applying AI and machine learning to structured data, but to apply it to the real world and unstructured data. Because 90 percent of the world’s data, and you can debate on the number, 80, 90, or more than 90 percent, is unstructured data. Whether that’s in text, images, video, audio.
That’s why I think our machine learning engineers and AI researchers were able to apply their knowledge and shift it toward a use case that was a bigger pain in the market for where we are today.
Chapter 11
Is It a Vision Problem or a Language Problem?
Dan puts it at roughly 80 percent vision, explains why OCR loses context that vision techniques recover, and describes how Landing AI reduces documents down to a point where a simple text model can finish the job.
Raja Iqbal: Do you look at the problems you’re solving in the document processing domain fundamentally as a vision problem, or a language problem, or both?
Dan Maloney: I’m going to say it’s a 70/30, 80/20 rule. From our point of view, it’s very much a vision problem. That’s really where we disrupted the dynamic.
When you looked at it classically as a language problem, that’s where a lot of work had to be done, extra coding, extra aspects. But when you blended the two together, when you brought the world of transformers, vision transformers, into this equation rather than just the classic OCR, that’s where a lot of the breakthroughs came. The agentic system, and how you can look at a thing multiple times and rework it, applying these techniques is really what solved a lot of the problems.
It’s always a blend of both. But when you make it at the end just a language problem, then it becomes much easier to solve.
Take OCR, for example. When it looks at something, and I’ll say generically this isn’t exactly 100 percent true all the time, but when you go up to the top left corner of a page and just start reading left to right: first of all, there might be languages that read right to left. You might be reading left to right, but it actually might be three columns, and it might go down, then up to the second column. So OCR can mess a lot of things up. It loses context, that awareness of so much that the document had. Those are vision problems.
Once we crack the code on vision problems, that’s where we can handle the really messy documents, the different modalities within the document, all those pieces. Once we can get it down to a text problem, a language problem, then it’s a much simpler problem. A lot of our effort, 80 percent, is more the real-world complex vision dynamics. But it’s not just one or the other, it’s both combined. I lean toward that classic 80/20 rule there.
Raja Iqbal: And do you think vision is harder than language, or language is harder than vision?
Dan Maloney: It’s not one harder than another. It’s, as you laid it out, when you break it down, what is the problem?
LLMs, and forget having multimodal, VLM, or other aspects in the equation there, that’s very difficult as well, what the LLMs have done. But they’ve done a great job solving a lot of the language problems. So when that becomes a potential challenge and we need to leverage it, it’s very easy for us to leverage. If we want to use one of the frontier models to solve a problem, a lot of money has already gone into that. They’ve solved a lot of those issues and we can use that effectively when it’s needed.
But so much of what we do, if you think about a document and you’re exploding it down to its fundamental blocks, even into its atoms, and building out a lot of those dynamics that we work on, when we break it and break it and break it down, once we do all that complexity, which is a lot of vision, structure, semantics, context, once we get it down to the simple part, then hopefully the majority of the time we can even solve it with a model that we’ve got that just handles text quite well. But there might be some generative dynamics in certain cases.
So when we break things down, I think of 80 percent vision. And then once we decide, are we transcribing it, we can handle that ourselves and we don’t need to involve a frontier lab to do that. Only when we need to solve a more complex problem where an LLM or some other model might come into play, we do that. And as those models get better, that even improves that small part of what we’re doing.
But the whole agentic system is able to tackle the problem, and the majority of that agentic system is much more focused around our visual AI techniques. That’s why it was such an easy next step for us to come into the world of document processing and disrupt the whole market.
Chapter 12
Earning Trust in Regulated Industries
Dan explains how Landing AI handles hallucination risk, and why serving banking and healthcare meant building SaaS, VPC, on-prem, and effectively air-gapped deployments.
Raja Iqbal: The kind of problems that you’re solving, having spent quite some time in AI and in industry, these systems can misfire. They can go wrong. So for regulated industries, how do you build systems in a way that earns the trust of not just the engineering team, but the compliance team, the regulators, and other people?
Dan Maloney: That’s a great question. We have to keep earning their trust every day, and we add more and more into the product, into the company itself, into our processes, the way we engage with them. But let me pull out a few hard tangibles.
Before I jump into the regulated industries, everybody still cares about trust. When I talk about autonomous and auditable, that auditable, that trust dynamic is a key part of the equation. Sometimes, depending on the use case, if you can guarantee that trust or auditability and that accuracy, they’re willing to pay a lot more for it, because they have to spend so much to make sure those things are accurate. So sometimes we can say, we can get you that level of trust or accuracy. But trust means much more beyond just accuracy. They’re even willing to pay a little bit more for it if we can solve it.
More broadly, one of the biggest challenges people often think about with LLMs is the potential for hallucination. So one of the first things we have to do around our accuracy is eliminate as much as we can. If we do leverage an LLM, we try to make it so simple and broken down for the LLM that when it receives something, it can be highly accurate and repetitive in that response. Every single time, it’s going to give that same response.
Early on when we launched, we had high accuracy, but we definitely saw deficiencies. And it really was across LLMs. It wasn’t one company. It wasn’t an Anthropic or a Google or an OpenAI. It was across the board. Once we started understanding accurately where they do well and consistently, and where they don’t, the places where they don’t is where we said, all right, we’re going to solve that more with a broader agentic system. We’re going to build our own special models, sometimes smaller, to solve those pieces, and we’re only going to hand over to the LLM what it can handle. Because sometimes giving them the whole document, it might do well at the front, maybe at the back, maybe it won’t do well in the middle. So we’ve solved that with a whole agentic system approach.
Secondly, that’s applicable regardless of regulated industry or not. Other key things that regulated industries need, which has also been a challenge: we have an offering delivered as SaaS, but we also had to deliver a VPC solution, so in a customer’s virtual private cloud, as well as an on-prem solution. Some companies go as far in the regulated industry that, for example in banking, their level 4 to level 7 data might not ever even be able to touch a commercial cloud, or something from one of the frontier labs.
So how do we solve that problem in our agentic system, where normally we might have a small component that goes out and uses an LLM? How do we do that internally? We had to either build our own models, which we’ve done in certain cases, and we’ve leveraged open source models and other aspects, and blend these together so it can in a sense be a complete air-gapped environment.
In our SaaS we might use some models. We go into a customer’s VPC, they might want to have other models. So we had to build that capability for model agnosticism. And then the on-prem is everything fully encapsulated, where it can never go out. In banking, sometimes healthcare, they might never have it even touch the cloud.
When you start delivering those deployments and can prove that you can still make it scalable, that it can be as accurate or more accurate, that it runs and can scale accordingly, that’s how you build trust with compliance, trust with their security team, trust with IT, trust with engineers, trust with the business units that are using it, to say, oh yeah, I tried it out in your SaaS and it’s just as good in my VPC or on-prem. That is not an easy thing to do.
You can choose to just stay in one world, but we’re really focused on developers in enterprises, and enterprises, and startups embedding us to sell their solutions into enterprises. So we need to manage that dynamic and that complexity for our end customers to earn that trust. It’s not just trust about the accuracy of the result, but handling data privacy, running in their environments and the complexities associated with that, so you can pass those compliance checks beyond just things like SOC 2 or zero data retention.
Chapter 13
Trust as Perception, and Visual Grounding
Raja points out that trust is often perception rather than fact. Dan explains visual grounding and confidence scoring, why Landing AI won’t claim 100 percent accuracy, and why customers sometimes discover errors in their own ground truth.
Raja Iqbal: As you mentioned trust, one more thought that crossed my mind is that in many ways, when it comes to the correctness of these systems, trust could be a perception, not based on some reality of facts. If a Tesla hits a pedestrian, it becomes news. But it happens all the time. So how do you overcome that challenge within an organisation? I’m just curious, how would you convince someone? Because all it takes is one bad extraction, one bad detection.
Dan Maloney: First of all, you’re absolutely right. We’ve tried to think about trust more broadly than just accuracy, confidence scoring. These are terms and techniques you might use to have somebody gain trust, but we’re literally re-looking at what elements we can give to an end user, whether that’s a developer, an analyst, a legal analyst, financial, whoever might be that user or consumer of it, that trust to make a decision to move forward.
So there are a few things. First of all, we can’t say that we’re 100 percent accurate, 100 percent right all the time. It’s a goal that we strive for, but it’ll be a goal that I believe we’ll always be striving for.
In our case, there are a couple of things we do to build that trust. One of the foundational pieces, which we tuned in the early years of our tech, is a technique we call visual grounding. Any element within our Agentic Document Extraction, any document that we process, a lot of times when you get down to text or elements within a table, we’re going to be able to show you visually where that answer came from, in a document, in a page, down to a chunk, down to a field level. Hey, I see Q4 FY26 revenue was $10 million. I’ll show you exactly where that number is and you can go validate it. Giving this human-level ability to go back and take a look at things is really valuable, really important, to help start building that trust.
Secondly, if you are not sure about an answer, giving the users a confidence score or some visual representation that says, hey, we might not be sure about this.
And think about that. If you’re looking at hundreds and thousands of pages, even just to get your ground truth, and I’m not talking about going through billions, if we can do a great job in helping them with what they already think is their ground truth. A lot of times companies take their ground truth and run it against us, and then they’re like, oh man, we actually had mistakes in our ground truth that we didn’t realise, that you called out. Which is a great thing to hear. So we’re even improving what they believe is 100 percent accurate, their golden set of that golden truth.
As we can do more and more at that trust level, to help with things like visual grounding, confidence scoring, validation dynamics, when you start layering all these elements together, that starts building up the trust. And when they start looking across hundreds or thousands of documents and they see that we’re performing 99 percent accurate or higher in the majority of their cases, trying it on their documents.
Anybody can go, and we’re guilty of it as well, customers have really asked us for benchmarks. We’ve done them. But benchmarks inherently can be overfitted. You can use a set of documents where you perform well. There are so many different ways to do it. All vendors try to do their best to be diverse. But sometimes you just have to try it on your own documents to see how it works, in our case in the realm of document processing. And we give that to them. They run it, see it, try it out in the real world, see how they were doing versus whatever they were doing in the past. That’s a level of us gaining their trust.
We’re adding more and more features into that. But really what we’re trying to do is build trust more broadly. It’s trust in the accuracy, trust in the security, the privacy. All that trust centre continues to be the most important part, especially when you’re selling into enterprises. That’s a layer in our product that we often talk about, and we’re never going to be done with it. We’re always going to be looking at new ways to gain the trust of customers so they can do really high scale.
Our ultimate goal is to run billions of documents through, and you never have to go look at it, and you just trust that our Agentic Document Extraction system is doing it 100 percent right. We’re not there yet, but I think we’re on the leading edge of all the companies out there.
Chapter 14
From Top Tier to SAP to Sapphire Ventures
Dan traces his career arc through sales, marketing, and development, a decade-long tenure at SAP after Top Tier was acquired, time in venture capital, and how a 2022 conversation with Andrew Ng brought him to Landing AI.
Raja Iqbal: Let’s shift a bit toward your journey as an entrepreneur, founding companies and being at other companies. You have built and sold companies, spent time in venture capital, and now you are again a CEO. Clearly you are more of a technical CEO, looking at your background. How have all of these roles, pure technical work, founder, CEO, venture capital, what did you learn from each of those roles that the other roles cannot teach you?
Dan Maloney: That’s a great question. My journey to CEO, as a multi-time CEO in this dynamic, I’ve met with tons of friends and peers who have been CEOs and founders, and I always find that everybody has a different path. Some can be fully technical, sales, marketing, product people.
My path was quite diverse. I’ve done sales, I’ve done marketing, I’ve literally done development, I’ve run major ecosystems. So I’ve gone through all functions. I wouldn’t say I’ve done accounting and finance, probably not my strength in the accounting arena. I ended up marrying somebody who’s into accounting, so she fills that for me, which is a benefit.
But really my path to CEO, or to these young companies that I’ve gone into, is really walking in everybody’s shoes. That to me was really helpful, and it really helps me connect the dots. In fact, sometimes when I was in functional roles inside different companies earlier in my years and I wasn’t CEO, I craved to connect all the dots. I was like, what role can I be in where I’m allowed to go across everything? Founder, CEO. In most roles you only have your incoming functions and your output. That’s why I almost naturally enjoyed being in this role, because I love having that strategic bigger-picture view and playing with all the different chess pieces on the board.
What was also interesting is that I loved tech early on. Back in fifth grade, sixth grade, my mom wanted me to get into computers. I liked playing computer games when I was young, on top of all my other activities, but I enjoyed that dynamic. I remember coming out and wanting to get into tech companies almost right away, and really I did right out of school.
One of my first companies was Top Tier Software, where I joined as probably a 24-year-old. We were all pretty young. I think Shai, my CEO, was maybe 28 years old, which I thought was really old at that time. I was like, he’s so wise, he’s four years older than me.
The company, three years in, got acquired into SAP. So I went from being in a startup into the number one software application company in the world at that time, and this is the early 2000s. All of a sudden that took my career in a completely different direction, dealing with the biggest enterprises in the world, one of the most powerful sales forces in the world, one of the most powerful ecosystems in the world at that time. I learned a lot about sales, marketing, go-to-market collectively, the ecosystem, and all those other aspects.
I thought we’d probably be there a year or two, but a lot of the Top Tier leadership got along quite well with SAP. A lot of the leadership went up into top positions. Shai even went onto the board. So we actually ended up having a decade-long tenure at SAP, a lot of us. Some are still there. That taught me a lot about the enterprise. And that gave me the connection into Sapphire Ventures, which used to be SAP Ventures.
While I was there, one of the jobs I enjoyed most was working with a lot of the ISVs in the ecosystem, all these startups, and watching them have success. When I was at Sapphire Ventures, being involved in all the different startups and seeing what they’re doing, I had that craving to get back into the startup world. I really enjoyed my time at SAP, it was extremely valuable, but I really wanted to get back.
So it was startup world, then enterprise company for a decade plus, and I think that grounding of the two together actually made me a better CEO and founder for my next companies. You mentioned a couple of them: Perspica, which we sold off to Cisco, and Zepl, which we sold off to DataRobot.
Andrew and I were both in the AI arena, and he was looking for a CEO and business person to run this company. After I wrapped up at my last company, I was thinking about what’s next and what direction to go. Do I go found a company from the ground up? Andrew and I had a good relationship, and we said, why don’t we do better together?
He’s the one who really sold me on it. I remember it was 2022. He said, it’s great what you’ve done in AI around structured data. The world of unstructured data and what’s going on right now is a massive revolution coming. And this was before we heard about all of it. Of course OpenAI was around in 2022 and others. But it was really that, when you think about the papers around transformers in 2017, vision transformers in 2020, out of the Google team, the Google Brain group, that led to where we are now. So we had a really good relationship, got along, and I was happy to join him on this journey.
Chapter 15
The Partnership with Andrew Ng Today
Dan describes how he and Andrew split the work, and tells the story of how SAP and Snowflake pointed Landing AI toward documents as their biggest pain point.
Raja Iqbal: And what does the partnership look like right now between you and Andrew?
Dan Maloney: Andrew is still founder of the company, still key in running all the deep tech for the company. He’s the one who works with all of our top MLEs and AI researchers. And he’s also moved into more of an executive chairman role inside the company.
All of the building of the business and the tech and the day-to-day pieces fall under me and the team. And then all of our new innovations fall with Andrew and I partnering up, along with the AI research team, who do an incredible job. That’s where Andrew really loves to spend his time, giving us input on the next-generation approaches.
In fact, we were sitting with Agentic Document Extraction probably two and a half years ago thinking about, can we really solve this problem? We were meeting with a couple of our partners, actually with SAP and with Snowflake, and we were showing off some of the next-gen tech we were doing. They were the ones who actually came to us. They said, this is amazing what you’re doing, we love that you’re tackling some of these challenges in visual AI. But specifically, our biggest pain point is documents.
What was surprising is that was a problem I was working on, among other things, back at SAP from 2001 to about 2012, 2013. I was shocked, when we double-clicked back into it, by how little it had moved in a decade and change. I think there wasn’t the tech to make the disruption at that time. So it was the right time and right place for us to come in. And again, some of our peers in the marketplace are now doing some great work on disrupting something that’s really been a struggle. The amount of innovation and resolution has been very minimal relative to the volume of documents and the complexities that have grown out there.
So I think finally we’re on the verge of being able to tackle something that’s been almost unsolvable for the past few decades.
Chapter 16
What Keeps Dan Up at Night
Dan names the two things always top of mind as CEO: the pace at which AI moves, and the continuous challenge of hiring the talent to keep up with it.
Raja Iqbal: Of all the problems that you have as the CEO, whether technical problems or otherwise, what is the problem that you find yourself thinking about the most?
Dan Maloney: That’s a good question. We’re very fortunate in this company to have deep AI knowledge and know-how, some of the best literally on the planet, especially in the realm of visual AI.
But the speed at which things are moving is still so fast that sometimes I think, wow, how do we keep up with this? How do I make sure my team’s empowered, that they have the budgets that they need? We’ve been talking about token maxing and these things early on. How do I make sure they have the latest and greatest? That has been a lot on our mind, and the transformation of AI and the impact of that on our company.
But I’ll say this: it’s something that, being here in Silicon Valley and being in almost like the bubble that we’re in, I worry about a lot. But then I go out and go around the rest of the world, the rest of the US, over to the East Coast, travelling over to Europe and Asia, and it’s okay. Even though it’s a continuous concern, how do we handle and stay on top of it, what I’m most excited about is that we’re trying to tackle that so our customers can just get the benefit from it.
It’s going to be very difficult when you don’t have the know-how and the personnel that we have in our company to really stay on top of it. But if we can provide that benefit, where we’re evaluating the latest and greatest models, looking at the new trends, and we can make that easier for our customers so they don’t need to worry about all the complexity, that’s a beautiful thing.
And if you don’t mind me putting the second part, it really is a challenge on a continuous basis to get the talent that we need in this market, in this arena. We’re very fortunate that we have Andrew and the team that we’ve got, who believe in the mission, where we’re focused on making the world’s documents computable and the impact for our customers.
Our engineers and our MLEs get so excited as we have more customers and the feedback comes in. You wouldn’t think that companies would be this excited about document processing. But when you’re solving problems that they’ve been having for 10, 20, 30 years, they literally get excited, and that’s really motivating for our team.
But I will say, we’re not a fast-growing employee company. We don’t believe we have to be a big company. But just continuing to get that talent that you need, who really understands what’s going on, that’s a challenge. Big companies are having it and small companies are having it. Getting that top talent is something that keeps me up at night.
Maybe in the future, AI is getting a lot of funding, maybe that settles down. If AI settles down and enterprises are paying so much, that might trickle back into the VC world. We haven’t seen that that much yet. But I would say AI, from the speed at which it’s moving, and the talent, are probably top of mind always as a CEO.
Chapter 17
Advice for Business Leaders: Don’t Boil the Ocean
Dan gives his practical guidance for leaders starting out with agents: pick one core use case, prove it end to end, and learn the retries, orchestration, and human-in-the-loop decisions that transfer everywhere else.
Raja Iqbal: Thanks. There’s a lot of excitement about agents automating entire workflows, and we talked about that in the earlier part of our conversation. So from where you sit, if a business leader is listening to this podcast, what is realistic, and how should they approach this transformation process? What is plausible and achievable as of today, what is plausible maybe next year, what is still a few years out, and what is not going to happen?
Dan Maloney: Sometimes I don’t know that I can predict the future about what will or will not happen exactly, but let me give you my view about where I’ve been and what I see going through this.
Even when I was back at SAP, we were building applications to drive a ton of automation throughout. There were still humans involved in transactions a lot of times, but a lot of the work started getting automated. Then we were looking at things like RPA, robotic process automation. You’re always looking at layering up more and more automation.
What is exciting about this is that yes, with agents and agentic systems, and with that agent-first approach, I do believe there’s a lot more in the workflows that truly can be completed end to end. Some of the challenges you might have, where something gets automated but you didn’t have checks and balances, or things being able to kick out to human in the loop, a lot more with AI now is a much more natural approach to automating workflows. So I’m actually very excited about it, from development of products all the way through to business workflows being fully automated.
The way I’d like to think about it is this. You have an opportunity right now to leverage AI agents, and the frameworks associated with that, to truly disrupt your business and automate a ton of different use cases and workflows in your company.
With that being said, if I’m getting started in this arena, this sounds almost like a cliché, but don’t try to boil the ocean. When you go everywhere, you really can’t measure that impact. When it’s too much, it’s like throwing sand into the ocean.
So what I’d recommend is grabbing use cases. Documents are a great area. It does not have to be that, if that’s not a major pain point. Finance, healthcare, legal, yes. There might be other areas where it’s not so critical. But grab a use case where you see it’s core to your business, it’s an area where you don’t believe you can throw unlimited humans and dollars at it, but it’s still going to be growing. If you can figure out those dynamics, put the effort into figuring out how to automate that, and really prove that in one example, with all the pieces that go along with it, you’ll figure out a lot more.
For example, in our world, some of the things that we figure out that customers don’t even think about. It’s like, oh yeah, what if we’re trying to process a document and something fails? We have to retry. How many times do you retry? Do I kick it off to another model? What are our agents doing? When do we kick it out to human in the loop?
So by going through that one or two use cases, you actually find out a lot about the metadata, the semantics around it, the orchestration, and then that can apply to other use cases downstream. But if you do too much at once, it’s overwhelming, it’s confusing. I feel like almost that’s what happened this past year, when everybody did token maxing, everybody went every direction. It was too hard for the enterprises, even very mature enterprises, and what they ended up with at the end of the year was a $100 million to $5 billion bill and not the results to show for it.
So I don’t think AI is so radically different from the way you tackled enterprise problems before, back in my days in the 2000s and 2010s. It’s just a technique where things can move faster, and that’s the great thing about it. You can tackle and automate a workflow that might have taken a month or months before. You can now do that in hours or days, which is pretty exciting. So you can test more, maybe even throw away more, because the testing time is so much quicker. That’s an exciting dynamic for business owners and business operators.
Chapter 18
The Snowflake Cortex Partnership
Dan explains why running natively inside Snowflake works like another deployment architecture, and how market capacity drawdown makes procurement easier for enterprise customers.
Raja Iqbal: You recently had this partnership with Snowflake Cortex. My understanding of that partnership is that it puts your visual AI inside their enterprise data infrastructure. So maybe the question is two parts. One, what does it change for your enterprise customers when your visual AI is right inside the data infrastructure? And do you see that as the future, that wherever there is enterprise data infrastructure, Landing AI is going to be plugged in? Is that a design pattern that should be replicated across the board?
Dan Maloney: Let me tell you my point of view on the partnership. I know the Snowflake team really well. Sridhar and Christian and the group are great, that’s their CEO and the head of product. The company is very partner-friendly. Their venture arm invested in us, very friendly from the point of view of investing in relevant companies they see with a nice overlap into their world.
One of the ways I think about Snowflake is not so different from how I sometimes think about a customer who might demand, hey, we need this in a VPC or on-prem. When a company sets up a data warehouse strategy, or even a broader data strategy, and Snowflake is at the core of that, it could be others, it could be Databricks, but let’s take Snowflake since you called that out.
Snowflake has literally set up an environment. It’s core for that company. They like that everything is protected inside. They set up the security within Snowflake. Even if they’re running things like Snowflake’s Cortex capabilities, they can host LLMs right within Snowflake, so it doesn’t necessarily go out to third parties, into the OpenAIs and Anthropics of the world. Everything is self-contained in that world.
So a lot of Snowflake customers are like, hey listen, I have this rich set of data in here. And by the way, my documents might be sitting on the outside. But if you’re telling me I can send a document into Landing AI, whether it’s natively inside Snowflake, or potentially outside of Snowflake, that’s okay, and then write the structured data into Snowflake. If you do that all within the four walls, let’s say natively inside Snowflake, a lot of companies find it easier to adopt, because they don’t have to worry about security concerns, just like if it’s in a VPC.
So think about Snowflake almost as another deployment architecture for us. If companies say, anything that runs inside Snowflake is easier to do business with, on the compliance, the security, we can trust that it’s not going to sneak outside the four walls, that’s a benefit.
And then there are other co-benefits. Snowflake has something called market capacity drawdown, which allows companies that have bought Snowflake capacity to use that capacity. So if they bought, I don’t know, $10 million worth of Snowflake, maybe they want to use $1 million of that to buy Landing AI’s Agentic Document Extraction. They can use that market capacity drawdown if they still have it available, and that makes the procurement part even easier. Again, you can do that as well with AWS. Azure has their own techniques.
I look at that as yet another win-win in a partnership. These companies have been using Snowflake for a long time and trust anything inside Snowflake or writing into Snowflake, and we made it easy for customers to make that decision if they’re interested.
But we need to be deployment flexible, because they may or may not have Snowflake. They may care that it runs inside Snowflake. Maybe they don’t care, they just want the data written into Snowflake. So we’ll make sure all those different options are available.
Chapter 19
What Excites Dan Most Over the Next 3 to 5 Years
Dan closes on the disruption he sees coming in document intelligence, and why freeing people from administrative work in medicine and government is what makes it worth doing.
Raja Iqbal: My last question. Looking three to five years ahead, what do you find most exciting? What are you most excited about in document processing and visual AI, whichever way they are headed?
Dan Maloney: In this broader realm of document intelligence, I think the thing that’s most exciting for me right now is the disruption that’s taking place.
Customers for the past several decades, as we’ve talked about, have struggled. They’ve had way too many humans still involved in this process. So much extra work, so many templates, so many fragile pieces. Because we look at it from an AI point of view and an agentic system approach, a lot of the complexity that just kept building and building, I fundamentally believe will be resolved, and will be resolved at a much cheaper price, with higher trust and higher velocity. That will dramatically save costs and unlock new revenue models and businesses that can be built because of it.
I think it’s fundamentally that impactful, because the data in those documents, in those images, in those other pieces, it’s like the world before discovering that you could use fossil fuels and oil, or harness the power of the sun. So I think we’ll make a big leap forward. That’s really what’s got me so excited, because I’ve seen the way it’s been done for the past 30, 40 years and it hasn’t been big incremental changes.
On top of that, there are the other bigger dynamics going on in AI. Once you start putting all of these building blocks together, these primitives, to form new ways for us to take care of businesses, enterprises, the Earth, medical challenges, new innovation, all of these things together.
Sometimes we’re talking over in the medical industry, and we’ve talked with nurses and doctors about how little time they get to spend with patients because they’re doing paperwork and administrative work. If we can solve that for the medical industry, for governments, and let the humans do what they were built to do, which is much higher level, much more important, much more strategic work.
A lot of times when you tackle a problem, you know what to do, but you get into it and then you get bogged down with administrative work. If we can eliminate that and let that human mind go, we could tackle the world. Those things combined are really what get me excited about the next three to five years and beyond.
Raja Iqbal: Dan, thank you so much for taking the time. Enjoyed the conversation.
Dan Maloney: Right back at you. I appreciate it, Raja. Thanks for the time today.
Raja Iqbal: Thank you so much.
Full episode on YouTube
Up next
More conversations
All episodes →
Agentic AI, sandboxing & enterprise security
Listen →EP 09João MouraFounder and CEO at CrewAI
Multi-agent systems, autonomous workflows & AI entrepreneurship
Listen →EP 08Emil EifremFounder and CEO at Neo4j