Transformation In Performance Audits: The Comptroller and Auditor General of India’s Technology Driven Audit Innovation
Author: Comptroller and Auditor General of India
Introduction
The expanding complexity of digital governance has fundamentally reshaped the audit landscape. Supreme Audit Institutions (SAIs) today confront voluminous financial records, and rapidly evolving programme architectures that strain conventional audit methods. Sample based scrutiny, and conventional data analysis methodologies are increasingly inadequate to detect systemic risks, hidden relationships, and governance failures embedded within interconnected digital systems.
The Comptroller and Auditor General of India is responding to this challenge through a structured technology driven audit innovation. Over recent years, our institution has integrated advanced analytical capabilities spanning network analysis, machine learning, optical character recognition (OCR), and AI-powered image analytics into its performance audit processes.
Case Study 1: Performance Audit of Welfare Schemes
India’s social protection architecture encompasses dozens of overlapping welfare programmes, from food security and housing to health insurance and rural employment, each maintaining independent beneficiary databases. While digital delivery through Direct Benefit Transfers (DBT) has substantially improved transparency, it has also created new challenges like exclusion of eligible households, silent duplication of benefits, and fragmented coverage across schemes. Conventional audits examined schemes in isolation, making cross-scheme systemic analysis almost impossible.
SAI India is conducting a data driven Beneficiary Schemes Audit covering seven major welfare schemes in seven States of India. The overall public expenditure involved in the seven schemes during 2023-24 was ₹4.14 billion. The audit adopted a household-centric, integrated analytical framework, treating beneficiary records across all seven schemes as interconnected nodes within a unified graph model. The combined size of the databases across all seven schemes is approximately 800 TB, wherein around 140 TB of data has been identified as relevant and required for the audit. Aadhaar1 hash values served as anonymised linkage identifiers, enabling the mapping of household-level benefit convergence and exclusion patterns without compromising privacy.
The analytical workflow encompassed dataset extraction and standardisation, cross-scheme household mapping, graph modelling and linked-entity analysis, anomaly detection, and risk-based prioritisation for field verification. The drone based validation to verify housing scheme records under the Pradhan Mantri Awas Yojana (PMAY, or an affordable housing mission by the Government of India) in selected locations revealed mismatch between stage of construction of house and release of funds.
Using big data analysis of the full population, we observed that approximately 3.6 million families remained excluded from benefits under major welfare schemes despite qualifying as eligible target households, a finding impossible to surface through conventional sampling and data analysis methods. This also highlighted the need for systemic reforms in targeting mechanisms and inter-scheme coordination, demonstrating how integrated analytics can deliver evidence based insights into welfare delivery effectiveness and equitable benefit reach.
Case Study 2: “Artifacial Intelligence” — AI Powered Image Analytics tool
Beneficiary verification in welfare audits has traditionally relied on structured data fields to detect anomalies. Yet a significant dimension of beneficiary identity has remained entirely unaudited i.e., the photographs stored alongside database records in virtually every major welfare scheme. Visual fraud like reused photographs enabling multiple registrations, blank or non-human images, group photographs obscuring individual identity represents a systematic vulnerability that structured data analysis cannot detect.
The Centre for Data Management and Analytics (CDMA) of SAI India has developed an AI based Image analytics toolkit called “Artifacial Intelligence” using open source Python libraries for image based beneficiary verification. This tool enables detection of duplicate images using perceptual hashing, flagging identical photographs used across multiple registrations with a 100% match threshold to minimise false positives, error image detection, identifying files with valid image extensions that contain no actual image data, face detection and count analysis.
Built on OpenCV, dlib, and face_recognition libraries with a PyQt6 graphical interface, the application requires no coding knowledge, runs on standard desktops, and produces outputs as structured CSV files or organised image directories. A Sankey2 diagram summarises analytics results visually for audit reporting.
The tool processed 238,000 photographs linked to 119,000 beneficiaries during a performance audit. Systemic flaws in the scheme’s IT infrastructure like image duplication and invalid entries that could have enabled misuse of welfare funds. The findings directly informed audit observations and policy recommendations for strengthening photo validation controls in welfare scheme IT systems.
Case Study 3: e-Procurement Audits through Machine Learning Algorithm
Public procurement remains one of the highest-risk domains for governance failures. Collusive bidding, cartelization, bid rotation, cover bidding, and the use of proxy vendors are persistent risks that conventional procurement audits struggle to identify. Tender-by-tender scrutiny, however thorough, cannot reveal the network-level patterns that characterise systemic collusion.
SAI India developed a relationship centric network analytics framework for e-procurement audits, integrating three complementary analytical techniques. The Graph-based Network Analysis modelled as interconnected nodes and edges, enabling visualisation of dense bidder clusters, dominant vendor networks, and anomalous procurement relationships at scale. Pattern Detection using Apriori Algorithm3 to identify recurring bidder combinations across multiple procurement events. Entity Resolution using fuzzy matching (which is a technique that is used to identify and link strings of data that may not be an exact match, but are likely to represent the same entity) where vendor names, addresses, and registration details were cross matched to detect duplicate vendor identities and concealed relationships between apparently independent bidders.
The framework was built entirely on open-source technologies like R programming and Python4 libraries ensuring both scalability and cost effectiveness. The methodology enabled full population analysis of procurement behaviour, detecting hidden vendor connections, collusive bidding patterns, and flag systemic risks invisible to conventional review. We also built institutional analytical capacity through standardised workflow templates, reducing the technical barrier for subsequent audit teams.
Case Study 4: OCR and AI-Assisted Document Analysis
A defining characteristic of public audit in large governance systems is the sheer volume of unstructured documentary evidence like scanned certificates, handwritten passbooks, invoices, assessment orders, etc. Manual scrutiny of this evidence is time-consuming, error-prone, and inherently limited in coverage.
SAI India deployed the Optical Character Recognition (OCR) technique across two distinct audit streams. It was deployed for welfare and document integrity audits which integrates bilingual OCR (Hindi and English) with Natural Language Processing (NLP) and computer vision to extract and analyse text from beneficiary certificates, passbooks, and related documents. A forensic layer detects tampering indicators like whitener marks, altered records, duplicate photographs and an audit fusion layer synthesises textual and visual evidence into explainable, AI-assisted audit conclusions. Similarly, in direct tax and financial statement audits, OCR extracts structured data from tax audit reports, invoices, computation sheets, and financial statements in PDF or scanned form. The extracted data is standardised, cross-verified across source documents, and analysed for anomalies and inconsistencies. Specific checks including invoice cut-off testing, cross verification between income tax returns and tax audit reports etc. are performed.
This has significantly reduced manual scrutiny workload, accelerated evidence processing, improved fraud indicator detection, and expanded audit coverage.
Practical Pathways for SAIs: Getting Started with Audit Innovation
SAI India’s experience across these four initiatives offers actionable lessons. The first stage is diagnostic and foundational where SAIs may begin by assessing the availability, quality, and accessibility of digital data held by audited entities. SAIs should conduct an honest assessment of in house analytical capacity: the skills available, the tools already in use, and the gaps that need bridging.
The second stage involves targeted experimentation. Rather than attempting institution wide transformation at once, SAIs are better served by identifying a specific audit area where digital data is available and the audit question is well-defined.
The third stage is methodological consolidation. Successful pilots must be codified into reusable frameworks: documented workflows, standardised scripts, audit templates, and quality assurance protocols that allow subsequent teams to replicate and adapt the methodology without rebuilding it from scratch. This is the stage at which innovation transitions from individual expertise to institutional capability.
The fourth and most critical stage is capacity and culture. SAIs must invest in structured training programmes that develop analytical literacy across audit cadres. This does not mean turning every auditor into a data scientist, but to ensure that analytical tools are understood, trusted, and used with professional judgement.
Conclusion
The digital transformation of governance systems demands a corresponding transformation in public sector auditing. SAI India’s experience demonstrates that emerging technologies like drones, image analytics, AI/ machine learning (ML) techniques can collectively shift audit capability from reactive transaction testing toward proactive, intelligence driven oversight.
As governance systems grow more complex and digitally interconnected, the SAIs that will deliver the greatest public value are those that invest today in the tools, techniques, and institutional culture required for tomorrow’s audit challenges. The transformation has begun, the imperative now is to accelerate and share it.
Footnotes
- Aadhar number is the unique 12-digit identification number issued to residents of India. ↩︎
- A Sankey diagram is a visual flow chart where the width of the arrows is directly proportional to the flow rate or quantity of data. ↩︎
- Unsupervised Machine Learning algorithm ↩︎
- Python libraries including NetworkX, igraph, and tidygraph ↩︎