ANALYTICS METHODOLOGY · VERSION 0.5
Is This A Good Car? Research Methodology
Is This A Good Car? uses the internal CarIQ analytics engine to support a transparent, data-driven used-vehicle research methodology. We do not currently publish a composite reliability score or rank model years.
CURRENT DATA
CarIQ currently uses official National Highway Traffic Safety Administration safety recall campaign records and owner complaint records. Source records are normalized into consistent internal types before descriptive statistics are calculated.
WHAT THE DATA CAN TELL US
The current data can describe how many complaint records were retrieved, which components those records identify, how concentrated component mentions are, how many records report structured outcomes, when reported incidents occurred, and the number and timing of recall campaigns.
WHAT THE DATA CANNOT TELL US
Current data cannot establish the percentage of vehicles that experience a problem, predict whether a particular vehicle will fail, confirm that every reported event resulted from a defect, or determine whether a used car is a good purchase. Raw complaint counts are not reliability rates.
DATA LIMITATIONS
- NHTSA-reported production is not the number of vehicles currently operating, and production-normalized complaint counts are not reliability or failure rates.
- Vehicle age, time on the road, model popularity, owner demographics, reporting behavior, publicity, and existing recalls can influence complaint volume.
- A complaint submitted to NHTSA is a consumer report, not automatically a verified vehicle defect.
- A recall campaign count does not measure reliability or describe whether an individual vehicle is affected or repaired.
- Different complaint outcomes should not be assumed equivalent. A set of minor reports is not analytically interchangeable with reports involving crashes, fires, injuries, or deaths.
- Missing source values remain missing. CarIQ does not estimate them without a defensible source.
HOW CARIQ ANALYZES COMPLAINTS
CarIQ preserves normalized complaint records separately from derived analytics. Combined NHTSA component labels are split transparently without merging distinct systems. Counts represent complaint records, category shares use records with component classifications, and the top-three concentration measures the share of categorized records mentioning at least one of the three most reported components.
Structured crash and fire values count complaint records reporting those outcomes. Injury and death values are totals reported across records. Mileage statistics use only finite values greater than zero; when mileage is unavailable, statistics remain null. Median and quartiles use sorted observations with linear interpolation.
PRODUCTION NORMALIZATION
Raw complaint totals are difficult to compare because vehicle populations differ. Where CarIQ can confidently match a catalog vehicle to official NHTSA Early Warning Reporting light-vehicle production records, it divides retrieved NHTSA owner complaint records by NHTSA-reported production and expresses the result per 100,000 vehicles produced.
EWR production is cumulative through a reporting period. CarIQ uses only the latest report for each manufacturer, make, model, model year, vehicle type, and platform subdivision, then combines distinct subdivisions of that same vehicle population. Quarterly values are never summed. Exact source identities are preferred; documented manual aliases may be used only after confirmation. Unmatched, conflicting, or ambiguous records produce no denominator.
Production volume is not equivalent to vehicles currently in operation, and complaints per 100K produced is not a failure rate. Reporting behavior, time on the road, publicity, recalls, and other factors affect complaint records. Vehicle age is not normalized because complaint reporting does not necessarily accumulate linearly.
MANUFACTURER COMMUNICATIONS
Manufacturer communications are technical information manufacturers submit to NHTSA. They may include service bulletins, repair instructions, service campaigns, warranty or policy extensions, software information, diagnostics, and other communications. CarIQ preserves the source summary and uses these records as descriptive evidence of conditions or procedures documented by a manufacturer.
Communication counts are not failure rates and do not show how frequently a condition occurs. Manufacturer documentation practices and level of detail vary, so raw counts should not be used alone to compare vehicle reliability. Component shares use only categorized communications, and a communication may mention more than one component.
NHTSA DEFECT INVESTIGATIONS
NHTSA defect investigations are formal agency activity evaluating potential motor-vehicle safety defects. They differ from consumer complaints, manufacturer service documentation, and formal recall campaigns. An investigation does not itself prove a defect, and no investigation does not establish that a vehicle is problem-free. CarIQ treats open or closed status as process information only.
CROSS-SOURCE ANALYSIS
CarIQ keeps complaints, manufacturer communications, recalls, and investigations separate because they describe different kinds of records. It identifies normalized component categories that appear across these sources without adding their counts or assigning weights.
Components are displayed first by the number of source categories containing records, then by complaint count, then alphabetically. Overlap describes available official documentation; it is not a failure rate, severity measure, reliability score, or proof of a defect.
DATA CONFIDENCE
Data Confidence measures evidence availability and completeness, not vehicle reliability. High requires all four descriptive sources, at least 25 complaints, at least 80% complaint component and incident-date completeness, a verified production denominator, and at least 50% mileage completeness. Moderate meets the four-source, sample, component, and date requirements but lacks production or mileage coverage. Limited requires at least three sources and at least five complaints but does not meet Moderate. Insufficient applies when fewer than five complaints are available, fewer than three sources are available, or complaint retrieval fails.
A successful source with zero records is available; a failed retrieval is unavailable. Recall or investigation counts do not increase confidence merely by being larger. Production is currently unavailable, so High is not currently expected.
CROSS-YEAR ANALYSIS
Model pages compare supported years chronologically from newest to oldest. Raw counts are descriptive. Older vehicles have had more time to accumulate records, and differing production volumes limit comparability. CarIQ does not rank model years, apply age normalization, or interpret count differences as improvement.
Production-normalized comparison remains unavailable until verified EWR production data can be obtained.
FUTURE METHODOLOGY DEVELOPMENT
CarIQ keeps raw source data, normalized records, identity matching, derived descriptive statistics, and any future composite methodology as separate layers. The methodology remains under development.
Analytics methodology version 0.5 is not a reliability scoring methodology.