The Data You Already Have That AI Can Use (and the Data You Think You Have but Don't)

Quick answer. You almost certainly already have data AI can use — a shared inbox, a spreadsheet of past jobs, or twelve months of invoices is often enough to start. What trips businesses up isn't a lack of data, it's assuming that data sitting in someone's head, on paper, or scattered across five disconnected tools counts as "data" an AI system can use. It doesn't, until it's pulled into one place.

Key takeaways

  • A messy spreadsheet with three years of history is more useful to an AI project than a clean, empty CRM.
  • A Melbourne electrical contracting business assumed it had "basically no data," but three years of quotes pulled from a shared Google Drive folder revealed it was quoting small jobs at a loss more often than anyone realised.
  • Twelve months of website enquiry forms is usually enough to build lead scoring or auto-routing.
  • The single highest-value task before an AI project is often just capturing one experienced staff member's tacit knowledge in a half-day interview.
  • Data spread across five disconnected tools isn't the blocker owners assume it is — untangling it into one shared view is often the actual first project, and it takes days, not months.

"We don't have enough data for AI" is one of the most common things we hear from Australian business owners in a first conversation. It's almost never true in the way they mean it.

What's actually true is narrower: most businesses have data, but it's not in a shape anything can use yet. That's a very different, and much more fixable, problem than "we don't have data." This post is about telling the two apart — in short, the data you already have that AI can use is broader than most owners assume.

The Data You Already Have That AI Can Use

It usually includes email inboxes, spreadsheets, accounting records, booking systems, call logs, and website forms. Before assuming you need to build something new, take stock of what already exists — most SMEs have several of these without realising they count.

Email and shared inboxes

Every customer enquiry, every quote request, every complaint your business has handled by email is a record of a real interaction. A few hundred emails is enough to see patterns: which questions come up most, how long responses typically take, what a good reply looks like versus a rushed one.

Spreadsheets, even messy ones

A job-tracking spreadsheet, a pricing sheet with a hundred manually-added rows, a roster built in Excel for the last two years — these are structured data, just not stored in a database. A messy spreadsheet with three years of history is more useful than a clean, empty CRM.

Accounting and invoicing records

Your ledger (Xero, MYOB, QuickBooks) already holds a detailed history of every transaction, every client, every payment timing pattern. That's enough to build things like cash flow visibility, automated invoice chasing tuned to each client's actual payment behaviour, or flagging unusual patterns.

Booking, job and scheduling systems

If you run a trade, clinic, salon or any appointment-based business, your booking platform holds a record of demand patterns, no-show rates, and client history that's often never looked at beyond the current week's calendar.

Call logs and voicemail transcripts

Many phone systems already log call duration, time of day, and sometimes transcripts. This is under-used data for understanding demand patterns and common questions, particularly for trades and home services businesses.

Website and form submissions

Every enquiry form, every quote request, every "book a call" submission is a labelled example of what a lead looks like. Twelve months of these is usually enough to build lead scoring or auto-routing.

Data source Usually assumed to be "not enough" What it's actually good for
Shared inbox / support email "Just a mess of old emails" Categorising queries, drafting responses, measuring response time
Job or booking spreadsheet "Just a spreadsheet, not a system" Demand forecasting, scheduling automation, pricing patterns
Accounting ledger "Just bookkeeping" Payment behaviour, cash flow forecasting, invoice chasing
Website enquiry forms "We don't track leads properly" Lead scoring, auto-routing, follow-up timing
Call logs "We don't record calls" Demand patterns, peak times, common questions

The data you think you have but don't

The data you think you have but often don't is anything that only exists as tacit knowledge, verbal handoffs, or records scattered across disconnected tools with no shared identifier. This is the less comfortable half of the conversation, and it's the one that actually matters for setting expectations before a build starts.

"It's all in my head" isn't data

If the pricing logic, the exception rules, or the way you triage urgent jobs exists only in an experienced staff member's judgement, that's not data yet — it's a process that hasn't been written down. The single highest-value task before any AI project is often just getting one person's tacit knowledge into a document. This usually takes a half-day interview, not a technical project.

Verbal handoffs

If information passes from the phone call to a sticky note to a verbal instruction to a technician, none of that is captured anywhere. It feels like the business "has" the information because a job gets done — but there's no record an automated system could learn from. This is common in trade businesses running on relationships and memory rather than software.

Data trapped in five disconnected tools

Having a CRM, a separate quoting tool, a separate accounting package and a separate scheduling app doesn't automatically mean you have integrated data — it often means you have four incomplete pictures of the same customer. Before assuming you're "data rich" because you use a lot of software, check whether any single record (a customer, a job) can actually be pulled together from all four without someone manually cross-referencing.

Historical data that's inconsistent or was entered inconsistently

A CRM where half the deals were logged properly and half were just "won" with no notes is technically data, but it teaches an AI system the wrong lesson if it's used uncritically. This doesn't mean it's useless — it means it needs a light clean-up pass first, which is normal and fast, not a six-month data project.

Data you're not allowed to use the way you assumed

If your data includes client information covered by privacy obligations (health records for an NDIS provider, for example, or financial data for a bookkeeping client), "we have the data" doesn't automatically mean "we can feed it into any AI tool." This is a scoping conversation to have early, not a reason to avoid AI — most of these cases have a straightforward, compliant path, but it needs to be designed in from the start.

A worked example

A Melbourne electrical contracting business came to Sketchli assuming they had "basically no data" because they didn't have a CRM. What they actually had: three years of quotes in a shared Google Drive folder, a job-tracking spreadsheet the office manager maintained daily, and an inbox with every customer email since the business started.

None of that was "a system." All of it was usable. We pulled three years of quotes and outcomes into one place and found the business was quoting jobs under a certain size at a loss more often than anyone realised — a pattern that had been sitting in plain sight in spreadsheets nobody had cross-referenced. That's a finding you get from organising existing data, not from buying new software.

Compare that to a business that assumes it's "data rich" because it runs four separate SaaS tools, but where a customer's history is split across all four with no shared identifier. Untangling that — even before any AI is involved — is often the actual first project.

How to check your own data before you commit a budget

You don't need a data audit consultant for this. A useful first pass takes half a day:

  1. List every place customer or job information currently lives — inboxes, spreadsheets, software, paper files, someone's memory.
  2. For each one, note how far back the history goes. A few months is workable for many processes; a few years is ideal but not required to start.
  3. Check whether records can be linked. Can you find the same customer's quote, job and invoice without manually searching three places?
  4. Flag anything sensitive — health, financial or identity information — so it's handled correctly from day one.
  5. Write down what exists only as tacit knowledge, and who holds it.

This maps directly onto how we run the discovery phase of an AI Readiness Audit — we're just describing the manual version here so you can do a first pass yourself.

FAQ

How much historical data do I actually need to start?

For most SME automation projects, a few months of consistent records is enough to start, and a year or more makes the results noticeably better. The exception is highly seasonal businesses, where at least one full cycle of data (often 12 months) gives a much more accurate picture. Don't wait to have years of perfect history — start with what exists and improve it as you go.

What if our data is spread across five different tools?

That's extremely common and not a blocker on its own — it's usually the actual first task. Connecting or consolidating a few key systems (even just exporting them into one shared view) often takes days, not months, and it's frequently where the first real insight comes from, before any AI automation is even switched on.

Is a spreadsheet good enough, or do we need a proper database?

A well-maintained spreadsheet is genuinely good enough to start most automation projects. It becomes a constraint only once a process needs real-time updates from multiple people simultaneously, at which point moving to a lightweight database is a sensible next step — but it's rarely the blocker owners assume it is on day one.

What about data protection or privacy rules?

If you handle sensitive information — health data, financial records, identity documents — this needs to be designed into the project from the start, not bolted on afterwards. It's a normal part of scoping a build properly, and most Australian SMEs already have reasonable safeguards in place from their existing software; the conversation is usually about extending those safeguards to the new automation, not starting from scratch. For anything security-specific, it's worth confirming current obligations with your own advisor or the relevant regulator.

Can AI help even if our data is genuinely messy?

Often, yes — a lot of early AI work is specifically about cleaning and structuring messy historical data well enough to be useful, which is a valuable outcome on its own even before automation is layered on top. The honest exception is data that's missing entirely rather than messy; you can clean up inconsistent records, but you can't clean up records that were never kept.

Should I fix our data before or after starting an AI project?

Usually alongside it, not before. Waiting for "perfect" data before starting is the single most common reason SMEs delay AI projects for years longer than necessary. A well-scoped first project — like the one described in What is AI transformation? — typically includes a light data clean-up as part of the build, rather than treating it as a separate prerequisite phase.

Next step

If you're not sure whether what you have is usable, that's exactly the question an audit answers. Try the free AI Readiness Check, read about our AI Readiness Audit, or book a free 30-minute call and bring a rough list of where your data currently lives.


Want to take your business to the next level with AI? Contact us or chat on WhatsApp.

Vish PrasadFounder & Product Lead, Sketchli

Sketchli designs, builds and launches AI-powered products and automations for first-time founders and growing Australian businesses, then stays until the numbers move.

Let's talk

Ready to find out what it would actually take?

Book a free 30-minute call. Tell us about your idea or your biggest bottleneck, and we'll tell you honestly whether AI can solve it, and exactly what it would cost. No pitch. No pressure.

or send us a note
Sent straight to Vish. Answered personally within one business day.
The Data AI Can Actually Use in Your SME | Sketchli