How to Run an AI Visibility Audit Without Expensive Tools

You do not need another opaque AI visibility score. This guide shows you how I use DataForSEO with Codex or Claude to track prompts, competitors, citations, AI referral traffic, conversions, and changes over time without paying for an expensive monitoring platform.

Most AI visibility tools sell you a score. The score looks reassuringly precise, but it rarely tells you which prompts produced the result, why a competitor appeared instead of you, which sources shaped the answer, or whether any of that visibility led to a useful visit.

That is why I do not treat an AI visibility checker as an audit. I use a controlled set of prompts, run it every two weeks, preserve the answers and sources, and connect the results to Google Analytics 4 and Google Search Console in Looker Studio. The system tells me what changed and gives me evidence for what to test next.

You do not need a costly monitoring platform to build it. My working stack is DataForSEO’s pay-as-you-go API plus Codex or Claude, a data store, Looker Studio, GA4, and GSC. A modest audit can cost a few dollars in API calls plus the AI subscription you may already use.

The point is not to replace judgment with AI. It is to use AI to remove repetitive work while keeping the method, definitions, and decisions under human control.

What Is an AI Visibility Audit?

An AI visibility audit is a structured review of when, where, and how a brand appears in generated answers. It measures mentions, recommendations, citations, competitors, source patterns, accuracy, and identifiable referral behavior across a defined set of prompts, platforms, locations, and dates. A useful audit preserves the underlying evidence instead of returning only one score.

That definition contains an important limitation: you are measuring a sample. You cannot test every question every buyer might ask, every variation a model might generate, or every personalized answer.

I call the result sampled prompt visibility, not total market share. If a brand appears in 18 of 30 tracked prompts, I can say it had a 60% mention rate in that controlled sample on that run. I cannot say it owns 60% of AI search.

The audit also separates four signals that are often blurred together:

  • A mention means the brand appears in the answer.
  • A recommendation means the system presents the brand as a suitable option.
  • A citation means the answer attributes or links supporting information to a source.
  • A referral visit means a person clicked from an identifiable AI assistant to the website.

A citation is not automatically a visit. A visit is not automatically a lead. A lead is not revenue. Keeping these layers separate makes the report less exciting at first glance and much more useful when you need to decide what to do.

What Should an AI Visibility Audit Measure?

An AI visibility audit should measure prompt-level presence, answer prominence, recommendations, citations, competitors, message accuracy, sentiment, referral sessions, landing pages, key events, and revenue where tracking allows it. It should also record the model, location, date, retrieval setting, and prompt version so that later comparisons are interpretable.

I organize the audit into four layers.

1. Sampled Prompt Visibility

For every prompt, I record whether the brand appeared, which competitors appeared, and where each brand appeared in the answer. Position can be useful, but it needs context: first in a ranked list means something different from the first brand casually named in an explanatory paragraph.

2. Mentions, Recommendations, and Citations

I store these separately. A company may be mentioned without being recommended, recommended without receiving a link, or cited as evidence without its brand being named prominently.

I also capture the cited URL and domain. Those sources reveal which owned pages AI systems retrieve and which third-party publishers influence the answer.

3. First-Party Website Behavior

GA4 shows whether identifiable AI referrals produced sessions, engaged sessions, key events, trials, purchases, or revenue. It also shows which landing pages received those visits.

Referral measurement is incomplete because many AI journeys do not produce a click, some referrers are lost, and someone may discover a company in ChatGPT before returning through branded search or direct traffic. Incomplete does not mean useless. It means you report what the data can and cannot prove.

4. Business Outcomes

Where tracking permits, I connect AI referrals to qualified forms, booked calls, sign-ups, trials, transactions, and revenue. A small channel that produces strong actions may matter more than a large visibility score that never reaches the website.

The AI visibility audit measurement system. A fixed prompt set produces answer-level evidence, which is connected to website behavior and business outcomes, then repeated every two weeks.

Why I Do Not Rely on a Free AI Visibility Checker

A free AI visibility checker can provide a quick snapshot, but it is not a substitute for a reproducible audit. You may not know which prompts, models, locations, repetitions, or scoring rules produced the number. Without that context, an AI visibility score can move while giving you no defensible explanation for why.

I will use a checker for orientation. I will not use one as the sole baseline for a client program.

The common weaknesses are practical:

  • The prompt set may be hidden or generated differently on every run.
  • Mentions, recommendations, and citations may be combined into one index.
  • The model version, country, language, or web-search setting may be unclear.
  • A single execution may be treated as a stable ranking.
  • Competitor and source evidence may not be exportable.
  • The score may not connect to GA4, GSC, leads, trials, or revenue.

AI answers are probabilistic. The same prompt can produce a different ordering or source set even when the website has not changed. That does not make measurement pointless; it makes documentation and repetition essential.

I want to be able to open any chart point and see the prompt, timestamp, platform, raw answer, mentioned brands, citations, and classification. If the tool cannot provide that chain of evidence, I treat its score as a lead for investigation rather than a conclusion.

The Low-Cost AI Visibility Audit Stack I Use

My low-cost stack combines DataForSEO for structured AI-search data, Codex or Claude for workflow automation and analysis, Looker Studio for reporting, and GA4 and GSC for first-party performance. DataForSEO is pay as you go, so a controlled prompt set can cost only a few dollars, although the total depends on endpoints, platforms, repetitions, and frequency.

ComponentWhat I Use It ForCost Context
DataForSEO APIPrompt responses, mentions, competitors, cited sources, pages, domains, and AI keyword dataPay per use; a focused audit may cost a few dollars
Codex or ClaudeBuild scripts, expand and classify prompts, normalize entities, compare runs, and flag anomaliesCommon individual subscription level is about $20 per month
Data storePreserve raw answers and normalized prompt-level rowsGoogle Sheets may be enough for a modest program
Looker StudioBlend visibility data with GA4 and GSCNo separate monitoring subscription required
GA4 and GSCReferral behavior, conversions, landing pages, queries, clicks, and impressionsFirst-party measurement tools

DataForSEO’s AI Optimization API supports LLM mention metrics, top mentioned pages and domains, source data, AI keyword volume, and live model-response or scraper workflows. Its current pricing is endpoint-specific, so I check the live pricing page before estimating a large run rather than publishing a universal cost per audit.

Codex or Claude handles the repetitive parts. I can give the agent an API schema, a prompt file, a scoring definition, and an output format, then have it build or update the collection and comparison workflow. I still review classifications, unexpected competitors, changed sources, and any recommendation that affects the website.

Need the audit, not another opaque score?

Phrase It’s GEO audit shows where your brand appears, who appears instead, which sources influence the answers, and what to change first. You get evidence and priorities rather than a dashboard with no decision attached.

Explore the GEO audit

Step 1: Build a Prompt Set Around Real Decisions

Start with questions buyers ask while defining a problem, comparing approaches, evaluating providers, and making a purchase decision. Use AI to expand and organize the set, but review every prompt manually. The final production list should be relevant, non-duplicative, and stable enough to run again without silently changing the test.

I build several prompt groups:

  1. Category prompts ask for companies, products, or services in the market.
  2. Problem prompts describe the pain without naming the solution.
  3. Comparison prompts ask how options differ or which option suits a condition.
  4. Recommendation prompts ask for suitable providers, tools, or approaches.
  5. Validation prompts ask about reputation, evidence, limitations, or alternatives.
  6. Branded prompts test how accurately the system describes the company.

The prompt should sound like something a person could plausibly ask. “Best B2B SaaS content agency for a technical product with a small internal team” gives you more diagnostic value than “top agency.”

I use Codex or Claude to find duplicates, label intent, identify missing decision stages, and create controlled variants. Then I remove generic prompts that would make almost any answer look relevant.

Keep a prompt ID that never changes. If you substantially rewrite a prompt, create a new version rather than overwriting the historical test.

Step 2: Define the Test Before You Run It

Document the prompt, platform, model, location, language, retrieval setting, run time, repetitions, and scoring rules before collecting results. This test contract stops normal model variation or an altered setup from being presented as a website improvement. It also makes the audit reproducible by another analyst.

At minimum, store these fields:

  • Prompt ID, text, category, intent, and version
  • Brand and known competitor entity names
  • Platform, model, market, and language
  • Whether live web retrieval was requested
  • Run date, timestamp, and repetition number
  • Raw response and raw citation/source data
  • Mention, recommendation, citation, position, sentiment, and accuracy labels
  • Notes about API errors, refusals, or ambiguous answers

Do not change ten variables at once. If you add a new platform, expand the prompt set, and alter the scoring rules in the same reporting period, the trend line no longer represents a comparable test.

Step 3: Collect AI Answers With DataForSEO

Use DataForSEO to retrieve structured LLM mention data or run controlled model prompts, then save the complete response before transforming it. Its endpoints can surface mentioned brands, top pages and domains, sources, search results, and cross-target metrics. That makes competitor and citation analysis much easier than copying answers into a spreadsheet by hand.

There are two related jobs here.

  • Historical discovery helps identify prompts, pages, domains, and competing brands already present in DataForSEO’s LLM-mentions dataset. I can investigate where a client or competitor appears and which sources recur around the topic.
  • Controlled monitoring runs my approved prompt set using consistent parameters. I preserve each response, normalize it into rows, and compare the same sample two weeks later.

Your exact endpoint mix will depend on the question. A brand-level audit may begin with target metrics and top-mentioned domains. A prompt experiment needs answer-level output and citations. A competitor-gap analysis may use several targets and group results by the same prompt or category.

DataForSEO returns raw JSON, which is useful because I am not locked into someone else’s dashboard. Codex or Claude can turn that JSON into a clean table with one row per prompt, run, brand, citation, or source. You can also use DataForSEO MCP.

Keep the raw response. Classification logic will evolve, and you may need to reprocess an old run without paying to collect it again.

Step 4: Use AI to Normalize and Review the Results

Use Codex or Claude to turn raw answers into consistent fields, compare entity variants, categorize sources, and identify changes between runs. Give the agent explicit definitions and require an evidence excerpt for every classification. Human review remains necessary for ambiguous brand names, implied recommendations, sentiment, and causal conclusions.

For example, the system can normalize “Phrase It,” “PhraseIt,” and phraseit.agency to one entity while retaining the exact text used in the answer. It can also separate owned pages, editorial publications, review sites, directories, social platforms, and competitor domains.

I ask the agent to flag rather than guess when:

  • A brand name could refer to more than one company.
  • A list has no meaningful order.
  • A source appears in search results but not in the final answer citations.
  • The answer praises a category but not the tracked business.
  • Sentiment depends on sarcasm or a complex qualification.
  • A change could result from model variation rather than a known intervention.

The agent accelerates review. It does not turn subjective labels into objective facts.

Step 5: Calculate Scores Without Hiding the Evidence

Calculate mention rate, recommendation rate, citation rate, competitor share within the sample, and accuracy issues separately. If you create an AI visibility index, publish its formula and keep the raw components visible. One weighted score may help executives scan a report, but it should never prevent an analyst from tracing the result back to individual prompts.

Useful formulas include:

  • Mention rate: prompts where the brand appeared / valid prompts tested
  • Recommendation rate: prompts where the brand was recommended / eligible recommendation prompts
  • Citation rate: prompts citing an owned URL / valid prompts tested
  • Prompt win rate: eligible prompts where the brand appeared and a named competitor did not / eligible prompts
  • Source concentration: citations from the most common source domain / all captured citations
  • Accuracy issue rate: branded prompts containing a material error / valid branded prompts

Always display the denominator. A jump from one mention to two is a 100% increase, but it still represents two prompts.

Position needs a written rule. If I score ordered recommendations, position one is clear. If the answer is prose with brands scattered across sections, I may record first textual occurrence and label it as prominence—not a universal rank.

Step 6: Turn Sources Into a Content and Outreach Map

Review recurring citations to learn which evidence AI systems use, which owned pages already contribute, and which third-party sources influence commercial answers. Prioritize publishers that are editorially relevant and already cover the category. A cited domain is an investigation lead, not automatic permission to pitch an unrelated link request.

I divide the source analysis into three buckets.

Owned Sources

Which pages from the audited website are cited? Look for patterns in topic, format, evidence, recency, author information, headings, and direct answers.

Competitor-Owned Sources

Which competitor pages are cited when the client is absent? The gap might involve a comparison page, an original dataset, clearer product information, a better explanation, or simply stronger existing demand.

Independent Sources

Which publications, directories, research pages, communities, or review sites recur? These may reveal legitimate digital PR, expert contribution, dataset, correction, or inclusion opportunities.

Outreach must have a reason. “Your page is cited by ChatGPT, please add us” is not a reason. New evidence, a missing comparison dimension, a verifiable correction, or a genuinely useful expert contribution is.

Step 7: Connect AI Visibility to GA4 in Looker Studio

Connect GA4 to Looker Studio and report AI referral sessions, engaged sessions, landing pages, key events, and revenue separately from prompt visibility. GA4 introduced a native AI Assistant channel in May 2026 for recognized assistants, but it excludes Google AI Overviews and AI Mode. A documented custom regex can provide additional control and continuity.

Google’s native classification is the easiest starting point. Use Session default channel group = AI Assistant for session reporting, or the relevant first-user dimension when acquisition is the question.

For a custom Looker Studio field based on the GA4 Session source dimension, I use a deliberately narrow allowlist:

REGEXP_MATCH(
  LOWER(Session source),
  “^(chatgpt\.com|chat\.openai\.com|perplexity\.ai|claude\.ai|gemini\.google\.com|copilot\.microsoft\.com|deepseek\.com|grok\.com|poe\.com|you\.com)$”
)

Name the field Known AI assistant source. Filter it to True, then build scorecards and tables for:

  • Sessions and engaged sessions
  • Engagement rate
  • Key events and session key-event rate
  • Total revenue when ecommerce tracking is valid
  • Session source and medium
  • Landing page plus query string

Do not use a loose pattern such as .*ai.*. It will match unrelated domains containing those letters. Maintain the allowlist, document the update date, and compare it with GA4’s native channel rather than assuming either captures everything.

Connect AI visibility to something the business can use

A mention is interesting. A cited landing page that produces trials, inquiries, or revenue is actionable. Phrase It builds reporting around both sides of that equation.

See Phrase It case studies

Step 8: Cross-Reference Google Search Console Data

Blend GSC with the prompt and GA4 data to understand the search context around cited pages. Compare queries, impressions, clicks, average position, landing pages, and branded demand with the topics represented in your prompt set. GSC does not report a clean ChatGPT ranking, but it helps explain whether the same pages are gaining broader discoverability.

I do not force a row-level join between a natural-language prompt and a Google query. They are different datasets.

Instead, I map both to shared fields such as topic cluster, funnel stage, landing page, market, or reporting period. That allows useful questions:

  • Are pages cited by AI also gaining relevant Google impressions?
  • Does a prompt-category improvement coincide with stronger visibility for the same topic?
  • Which high-impression pages never appear as AI sources?
  • Did branded search rise after the company began appearing in recommendations?
  • Which landing pages receive AI referrals but perform poorly after the click?

The dashboard should encourage investigation, not manufacture attribution.

Step 9: Repeat the Audit Every Two Weeks

Run the same controlled prompt set every two weeks, store each raw response, and compare mentions, recommendations, citations, competitors, sources, and accuracy. Add website and outreach annotations to the timeline. A biweekly cadence is frequent enough to catch movement without mistaking daily model variation for a strategic trend.

My cycle is straightforward:

  1. Rerun the frozen production prompt set.
  2. Preserve raw answers, model details, citations, and errors.
  3. Normalize entities and classifications with the same rules.
  4. Review changes, new competitors, lost citations, and new source domains.
  5. Compare AI referrals, landing pages, key events, and revenue.
  6. Review GSC performance for the relevant pages and topics.
  7. Add annotations for content, technical, product, and authority changes.
  8. Select the next hypothesis and define how it will be evaluated.

I keep experimental prompts outside the main trend until they have a stable definition. Otherwise, adding ten easy branded prompts could make visibility appear to improve overnight without the brand becoming more discoverable.

Repeated tests also need restraint. A changed answer after one edit is an observation. Consistent movement across prompts, runs, models, and supporting first-party signals is stronger evidence, but it still may not isolate causation completely.

Step 10: Analyze What Gets Cited, Form a Hypothesis, and Test It

Compare cited and uncited content to identify a plausible reason for inclusion, then change one meaningful variable and monitor several cycles. Useful hypotheses involve original evidence, clearer answers, better entity information, updated facts, technical access, or independent corroboration. Do not rewrite a whole site around one screenshot from one model.

When a page is cited, I inspect:

  • Whether it supplies original data, examples, definitions, or comparisons
  • Whether important claims are explicit and easy to quote accurately
  • Whether the author, company, product, date, and market are unambiguous
  • Whether the page is accessible without scripts, login walls, or crawler blocks
  • Whether third-party sources corroborate the entity or claim
  • Whether the cited passage directly answers the prompt

Then I compare it with pages that were eligible but ignored.

An experiment might add a transparent methodology to an existing research page, clarify pricing conditions, publish a comparison the market lacks, correct inconsistent entity descriptions, or earn coverage from an already influential industry source. I annotate the implementation date and watch the same prompt sample.

AI behavior can change independently. That is why I look for repeated directional evidence instead of declaring a win after the first favorable response.

What Phrase It’s AI Referral Data Shows

For one client, AI referral sessions reached 1,458, up 21.4% from the comparison period. Key events increased 90.5% to 80, and free trials rose 117.6% to 37.

The same reporting showed 939 engaged sessions. ChatGPT accounted for 1,246 of the 1,458 AI sessions, 69 of the 80 key events, and 35 of the 37 free trials. Perplexity and Gemini produced smaller but measurable conversion activity, while Claude and Microsoft Copilot contributed referral traffic.

Those numbers matter because visibility alone would have missed the most interesting part. Sessions grew, but key events and trials grew much faster.

You can review Phrase It’s published evidence in the work and case-study library. The broader LLM SEO guide explains how technical access, content, authority, citations, and measurement fit together.

When Is a DIY AI Visibility Audit Enough?

A DIY audit is usually enough when you have one market, a modest prompt set, access to GA4 and GSC, and someone comfortable reviewing API output and ambiguous AI answers. Professional help becomes more useful when the scope spans markets, models, competitors, source outreach, technical problems, implementation priorities, and stakeholder reporting.

The API is not the difficult part. The difficult part is deciding what counts, preserving comparability, interpreting conflicting signals, and turning the evidence into work that might change the outcome.

A DIY process is a good fit when:

  • You can define one commercially relevant prompt sample.
  • You can maintain the collection and reporting workflow.
  • You have enough first-party traffic or events to interpret cautiously.
  • You are willing to review classifications and limitations.
  • You can implement and annotate experiments.

An audit needs deeper support when different teams own analytics, SEO, content, PR, and development, or when executives need one defensible reporting model across regions. Phrase It’s GEO audit service is built for that situation: find where the brand appears, who appears instead, which sources shape answers, and what should change first.

An AI Visibility Audit Should Produce Decisions, Not Just a Score

The useful output of an AI visibility audit is not a colorful number. It is a chain of evidence: which buyer prompts you tested, what the systems answered, where your brand and competitors appeared, which sources shaped the answers, which pages earned visits, what those visitors did, and what changed after your intervention.

Build the small version first. Freeze a commercially relevant prompt set, collect it through DataForSEO, use Codex or Claude to normalize the work, and connect the result to GA4 and GSC in Looker Studio. Run it again in two weeks.

You will still have uncertainty. But you will know exactly what the uncertainty depends on, and you will have a much better basis for the next content, technical, or authority decision.

Find out what AI search is actually using

Bring your website, competitors, and priority market. I will show you how I would scope the prompts, sources, measurement, and fixes behind a defensible AI visibility audit.

Book a 30-minute strategy call

Frequently Asked Questions About AI Visibility Audits

How Much Does an AI Visibility Audit Cost?

Our targeted GEO audit costs from $1200, $2000 and $4000 depending on the scope. The cost depends on the number of prompts, platforms, markets, repetitions, API endpoints, analysis depth, and whether implementation recommendations are included. My low-cost DIY setup can use only a few dollars in DataForSEO calls for a controlled run, plus a roughly $20 monthly Codex or Claude subscription. Large monitoring programs cost more.

Can I Run an AI Visibility Audit for Free?

You can manually test a small number of prompts and use free checkers, but the result will be difficult to reproduce at scale. A small pay-as-you-go API budget creates a stronger audit because you can preserve structured responses, competitors, citations, and run settings without buying a full monitoring platform.

How Often Should I Check AI Visibility?

I run the controlled prompt set every two weeks. Daily checking creates noise for most strategy programs, while quarterly checking can hide when a citation, competitor, or description changed. Use the same cadence and test conditions long enough to distinguish persistent movement from normal answer variation.

What Is the Difference Between an AI Visibility Audit and an AI Visibility Checker?

An AI visibility checker usually returns a quick snapshot or score. An audit documents the prompts and test conditions, preserves raw answers, separates mentions from citations and recommendations, examines competitors and sources, connects results to first-party analytics, and produces prioritized experiments.

Can GA4 Track Traffic From ChatGPT and Other AI Assistants?

Yes, when the assistant passes an identifiable referrer. GA4’s native AI Assistant channel recognizes sources such as ChatGPT, Gemini, DeepSeek, Copilot, and Grok, but Google says the channel excludes AI Overviews and AI Mode. Referrer loss and zero-click journeys mean GA4 cannot measure every AI-assisted discovery.

Can an Audit Show Why Competitors Appear Instead of My Brand?

It can identify evidence for a hypothesis by comparing answer wording, competitors, citations, source domains, and the content each source provides. It cannot reveal a model’s complete internal reasoning. Use the patterns to design a controlled content, technical, entity, or authority experiment and monitor whether the result persists.

Is an AI Visibility Score the Same as Share of Voice?

Not automatically. A visibility score may use a proprietary weighted formula, while share of voice requires a defined universe and denominator. For a controlled audit, report the percentage of eligible tracked prompts containing the brand and call it sampled visibility unless you have defensible prompt-volume weighting.

Tell us what good looks like for you.

Share your site, goals, and what is getting in the way. We'll reply with the clearest next step, whether that is working together or fixing something first.

We reply within one working day. Always a real human.
Scroll to Top