
MCA Data Is a Good Starting Point, and Only a Starting Point
Twenty audited filings, one tidy distribution, and a good part of the spread has nothing to do with how those companies perform.
Summary
- Indian filing data is better raw material than most markets provide, and it settles financials only. Unit economics, cycle times, customer metrics and capability maturity are never in it.
- Lease accounting is the largest invisible distortion. An Ind AS 116 adopter carries lease cost below EBITDA and a non-adopter carries it above, so identical economics produce different margins.
- For a smaller company the accounts are often prepared for tax rather than to represent economics. Saying a credible peer set cannot be built is a better answer than publishing one anyway.
India's filing regime gives analysts something most markets do not. Every registered company files. The filings are audited, prepared to a common format, and available going back years. Vendors like Probe42 and Tofler structure them into usable datasets, and CMIE and Capitaline extend the history for larger entities.
Compare that to most markets, where private company financials are simply unavailable and peer analysis is confined to listed companies. In India you can build a peer set of twenty mid-market private companies and read their audited accounts. That is a real asset, and it is the foundation of every benchmarking exercise we run.
What it is not is a finished dataset. Pull the filings of twenty companies in the same sector, calculate margin for each, and you will produce a distribution that looks authoritative. A meaningful part of the spread in it will have nothing to do with how those companies actually perform.
The work between the filing and the benchmark is specific, known, and largely invisible in the output, which is exactly why it gets skipped. This piece sets out what that work involves.
This is the fourth piece in a series on benchmarking. The earlier ones covered how to frame a benchmarking exercise, the metrics that carry a diagnosis, and when revenue multiples apply.
What MCA data settles, and what it does not
It settles a great deal. Revenue, cost structure, balance sheet, borrowings, charges, shareholding, directors, related-party disclosures and auditor observations, for essentially every company in the country, on a consistent statutory basis, year after year.
What it does not settle is comparability. The filings are prepared to a common standard. Within that standard, companies make different presentation choices, adopt standards at different times, structure themselves differently, and file to different year-ends. None of those differences are errors. All of them distort a naive comparison.
There is also a coverage limit worth stating plainly. Filings carry financials only. , functional ratios, cycle times, customer metrics and capability maturity are not in them and never will be. Those come from the company's own systems or from primary collection.
Get the peer set right before any number is compared
Before any number is compared, the peer set has to be right, and this is where automated screening fails first.
NIC codes are self-declared at incorporation and rarely updated. A services business and a trading business can sit under the same code. Diversified groups file under whatever the holding entity was originally registered as, which may describe a business the group exited a decade ago.
Screening on classification alone therefore produces a distribution that measures classification error as much as performance. The reliable method is to cast a deliberately wide net: codes, keyword searches on company objects, known competitors, and peers named in listed players' annual reports. Then open the website and the latest filing for every company that survives, and confirm what it actually does.
It is slow and it does not delegate well. It is also where the credibility of everything downstream comes from, and it is the part of market intelligence work that never appears in the deliverable.
Why lease accounting quietly reorders a peer ranking
Lease is the single biggest source of false variance in Indian EBITDA comparison, and almost nobody adjusts for it.
A company that has adopted Ind AS 116 carries lease costs in depreciation and finance cost. A company that has not carries the same economic cost as rent, inside operating expenses. Same business, same leases, same economics, and a materially different EBITDA margin.
For retail, healthcare, warehousing, hospitality and any business with a significant leased footprint, this alone can reorder a peer ranking. Adoption status has to be confirmed company by company, and the adjustment made or the divergence flagged.
Rebuild EBITDA rather than reading it
EBITDA is not a defined term under Indian GAAP or Ind AS. Every data vendor calculates it differently, and companies present it differently in their own disclosures. Taking a reported or vendor-calculated figure into a peer comparison imports that inconsistency directly into the analysis.
The rebuild runs from the schedules, applied identically across every entity in the set.
- Other income excluded. Treasury income in a cash-rich company otherwise flatters operating margin, sometimes substantially in businesses sitting on fundraise proceeds.
- Exceptional items treated consistently. Consistency across the set matters more than the specific treatment chosen.
- ESOP charges handled the same way for everyone. A funded company expensing ESOPs heavily is not comparable at the EBITDA line with a promoter-owned company that has none.
- Add-backs limited and logged. Non-recurring items and one-off provisions, with every adjustment recorded against its reason.
Revenue is not always revenue
Several presentation choices change the revenue line by enough to make margin comparison meaningless.
Gross versus net presentation in marketplace, trading and agency models. One company books the full transaction value, another books only its commission. Both presentations can be correct under the standards, and their margins are incomparable by an order of magnitude.
Related-party revenue. The RPT schedule needs reading. Material inter-group revenue should be shown separately rather than netted silently. A group entity with a large share of revenue from affiliates has a different business from what the top line suggests.
Other income inside revenue. Some companies present it within revenue from operations. It does not belong in an operating comparison.
Which entity, which year, and how far behind?
Standalone versus consolidated cannot be mixed within a peer set, and the choice has to be deliberate: consolidated where operating subsidiaries exist, standalone only where they do not. Where an entity's basis changed across the years being compared, that needs flagging rather than smoothing.
Holding structures require tracing to the operating entity. Indian groups routinely split revenue, costs and assets across a manufacturing entity, a trading entity and a property entity, and any single filing describes a fragment. Either recombine or exclude.
Fiscal year variance. Most companies report to 31 March, but not all, and a peer whose year-end differs by more than a quarter is describing a different economic period.
Lag. Private filings run well behind the current year. When client data is current and peer data is not, the comparison spans different periods. The report should say which period each side reflects rather than presenting them as contemporaneous.
When filings are not enough to benchmark an MSME
For smaller companies, the normalisation above is necessary but not sufficient, because the accounts are frequently prepared for tax purposes rather than to represent economics.
Promoter remuneration set for tax reasons rather than market rates. Family expenses inside the P&L. Rent paid to promoter-owned entities at non-arm's-length rates. Abridged filings that disclose too little to reconstruct anything. The distortion runs in both directions, and it is not a data quality problem so much as a different purpose of preparation.
The practical consequence is that peer benchmarking sometimes cannot be done credibly for an MSME. The honest sequence is to normalise to real economics first. Then benchmark internally and historically, where the data is at least consistently prepared. Then lead with maturity benchmarking, which needs no external comparables at all. It is the same threshold we write about in why a founder's gut stops working past ₹20 crore.
Saying a credible peer set cannot be built is a better answer than publishing a distribution assembled from a handful of unreliable filings.
Three habits that make a benchmark checkable
Three habits, none of them complicated, all of them frequently skipped.
Publish the definitions. Every metric in a benchmarking output should carry its formula. Where definitions are visible, a reader can check the work; where they are not, they can only take it on trust. It is the same principle as investors funding their confidence in your numbers, rather than the numbers themselves.
Disclose the judgments. Where an adjustment involves discretion, meaning which add-backs and how exceptional items were treated, state it. A reader who disagrees with a specific judgment can still use the analysis if they can see it.
State the sample. Percentiles need a meaningful number of comparable entities. Below that, ranges with the count disclosed. Below that again, individual comparisons only.
The underlying point is straightforward. Indian filing data is better raw material than most markets provide, and it produces reliable benchmarks once the known distortions are handled. What it does not do is produce them automatically. The difference between a benchmark that holds up and one that quietly falls apart is entirely in work that never appears in the final chart.
How useful was this article?
One tap. It tells us what to write more of.
The next one
Get what we publish next, by email.
Working notes on raising, borrowing, protecting, growing and structuring capital in India. One email a week at most, and you can leave any time.
We use your address only to send this. See our privacy policy.
We store your address to send you these emails and nothing else. See our privacy policy.
Related reading



