Methodology

Transparent documentation of how we collect, clean, and aggregate search data from 15+ directory and mapping platforms.

🔍

Data Collection

Multi-source aggregation from search engines, directories, and SEO platforms

1

Google Keyword Planner API

Primary search volume source. We query 10,000+ local service keywords across 200+ US metro areas monthly. Volume ranges are averaged over 12 months to smooth seasonality.

  • Official Google Ads API
  • 12-month rolling average
  • Metro-level granularity
  • Exact match + close variants
2

Google Maps Platform (Places API)

Business listing data, review counts, ratings, photos, hours, and popular times. Used to validate directory coverage and enrich category data.

  • Place details
  • Review analysis
  • Popular times
  • Business status verification
3

Yelp Fusion API

Yelp's business directory data including review content, categories, attributes, and search rankings. Strong for restaurants, home services, and local retail.

  • Review text analysis
  • Category hierarchy
  • Attribute extraction
  • Check-in data
4

Vertical Directory APIs & Partnerships

Direct data from Zocdoc (healthcare), Angi (home services), OpenTable (restaurants), Houzz (renovation), Avvo (legal), Thumbtack (general services). Each has unique search behavior patterns.

  • Appointment availability
  • Lead volume estimates
  • Booking conversion rates
  • Specialty taxonomies
5

SEO Intelligence Platforms (SEMrush, Ahrefs)

Third-party keyword difficulty, SERP features, competitor analysis, traffic estimates, and backlink profiles. Used for validation and gap-filling.

  • Keyword difficulty (KD)
  • SERP feature tracking
  • Competitor gap analysis
  • Traffic estimation
⚙️

Data Processing & Normalization

Cleaning, deduplication, and standardization across sources

Keyword Normalization

Map variant spellings, synonyms, and regional terms to canonical keywords. 'HVAC repair', 'air conditioning repair', 'AC fix' → unified 'HVAC repair' cluster.

  • Fuzzy matching + manual review
  • Regional dialect handling
  • Brand name filtering
  • Intent classification

Volume Aggregation

Weighted combination of sources. Google KWP = 50% weight, vertical directories = 30%, SEO tools = 20%. Outlier detection removes anomalous spikes.

  • Weighted averaging
  • Outlier detection (3σ)
  • Seasonal adjustment
  • Confidence intervals

Geographic Mapping

Metro-level volumes mapped to categories. Population-weighted allocation for national estimates. Rural areas use state-level fallback with density adjustments.

  • MSA boundaries (Census)
  • Population weighting
  • Density corrections
  • Cross-border metro handling

Category Taxonomy

Unified 15-category taxonomy with 200+ subcategories. Maps each source's native categories to our standard. Manual curation for edge cases.

  • 15 primary categories
  • 200+ subcategories
  • Source-to-standard mapping
  • Quarterly taxonomy review

Quality Assurance

Validation processes ensuring data accuracy

Volume Cross-Validation

AutomatedManual Review

Compare Google KWP volumes against SEMrush/Ahrefs estimates. Flag discrepancies >50% for manual review.

Category Coverage Audit

Automated

Verify each category has ≥3 data sources. Categories with single-source dependency flagged for expansion.

Seasonality Sanity Check

AutomatedManual Review

Detect unusual spikes/drops vs historical patterns. HVAC in July, tax prep in April — validate expected seasonality.

Geographic Consistency

Automated

Metro volumes should correlate with population × service density. Outlier metros investigated for data quality issues.

Vertical Directory Reconciliation

Manual Review

Compare vertical directory lead volumes against Google search volumes. Large gaps indicate tracking or market changes.

New Category Detection

Manual Review

Quarterly keyword research for emerging service categories (e.g., "EV charger installation", "AI consulting").

⚠️

Known Limitations

Transparency about what our data cannot tell you

  • 1

    Google KWP Volume Buckets

    Google reports volumes in ranges (100-1K, 1K-10K, etc.), not exact numbers. We use midpoint estimates with confidence intervals.

  • 2

    No Click-Through Data

    Search volume ≠ clicks. Zero-click searches (Maps pack, AI overviews) mean volume overstates traffic potential. We don't have CTR data.

  • 3

    Vertical Directory Opacity

    Most vertical directories (Zocdoc, Angi, Houzz) don't publish search volumes. We infer from lead volumes, partner disclosures, and third-party estimates.

  • 4

    Geographic Granularity Limits

    Metro-level is finest granularity from Google KWP. Neighborhood-level search behavior inferred from Maps data, not direct volume measurement.

  • 5

    Brand vs. Generic Queries

    Branded searches ("Roto-Rooter near me") inflate category volumes. We filter known brands but miss emerging/local brands.

  • 6

    Voice Search Attribution

    Voice queries routed through same KWP but with different intent patterns. Cannot fully separate voice vs. typed without Google Search Console data.

  • 7

    International Coverage

    Currently US-only. Canada/UK/AU have different directory ecosystems and search behaviors. Expansion planned for 2025.

📅

Update Cadence

Monthly refresh cycle with annual methodology review

Monthly (1st-5th)

Data Refresh

Pull fresh volumes from all APIs, re-run normalization, update category pages.

Quarterly

Taxonomy Review

Add emerging categories, merge declining ones, update subcategory mappings based on search trend analysis.

Annually

Methodology Audit

Full source re-evaluation, weight recalibration, new source integration, limitation reassessment.

Questions About Our Data?

We're transparent because trust matters. If you have questions about sources, methodology, or limitations, reach out.