# Differences and Integration of ESG Ratings
## Introduction: The Alphabet Soup That Decides Capital Flow
If you’ve spent any time in financial circles over the past five years, you’ve likely heard the acronym ESG thrown around like confetti. Environmental, Social, and Governance—three pillars that supposedly measure how “responsible” a company truly is. But here’s the kicker: ask five different rating agencies to score the same company, and you’ll get five wildly different numbers. Sometimes even rankings that contradict each other. It’s a paradox that keeps CFOs up at night and makes portfolio managers question their own risk models.
I remember sitting in a strategy meeting back in 2021, when one of our clients—a mid-sized renewable energy firm—showed us their ESG scores from three major providers. One agency gave them a stellar 82/100. Another scored them a lukewarm 54. The third flat-out refused to issue a rating due to “insufficient data.” Same company. Same year. Same public disclosures. Different worlds. That moment crystallized something I had been feeling for a while: **the ESG rating ecosystem is less a unified framework and more a collection of competing philosophies dressed up as science.**
This article isn’t just about complaining, though. It’s about understanding *why* these differences exist, and more importantly, *how* we can move toward integration without losing the richness that diversity brings. Because let’s be honest—a single monolithic ESG score would be convenient, but it would also be dangerously simplistic. The goal isn’t to make everyone agree. The goal is to make the disagreement *useful*.
Let’s dig into the meat of it.
## The Measurement Mismatch: What Are We Even Scoring?
The first and most fundamental source of divergence lies in the *definition* of the metrics themselves. When MSCI says “governance,” do they mean board diversity? Executive pay ratios? Anti-corruption policies? Or all of the above, weighted differently? The answer is yes, but the weights are proprietary secrets locked behind paywalls.
Take carbon emissions, the most “objective” of ESG metrics. Even here, agencies disagree on scope. Scope 1 (direct emissions) is easy. Scope 2 (energy purchased) is moderately tricky. Scope 3 (supply chain and product use) is a swamp of estimates and assumptions. Some agencies include Scope 3 with a 30% weight; others ignore it entirely. A company with a sprawling global supply chain could look like a climate hero under one methodology and a villain under another—same facts, different lenses.
I’ve seen this play out in real time. At BRAIN TECHNOLOGY LIMITED, we once ran a comparative analysis for a logistics client. One agency penalized them heavily for not having a net-zero target for their shipping fleet. Another gave them credit for using biofuel blends, even though the first agency didn’t even consider biofuel a valid mitigation measure. Neither was “wrong.” They were just answering different questions with the same data.
The core issue is that ESG ratings are not measurements; they are interpretations. And interpretations require assumptions. Assumptions about materiality, about time horizons, about stakeholder priorities. Until we have a universally accepted taxonomy of what counts as “good” ESG performance—which, frankly, may be impossible—we will keep seeing these mismatches.
## Weighting Woes: How Materiality Depends on Whom You Ask
Let’s talk about weights. Even if two agencies agree on the metrics, they rarely agree on the *importance* of each metric. Is water usage more material than workforce turnover? Depends on whether you’re a beverage company or a software firm. The concept of “materiality” is supposed to address this, but different agencies approach it differently.
SASB (Sustainability Accounting Standards Board) has built an entire industry around industry-specific materiality maps. But even within SASB’s framework, you have to decide whether environmental factors get a 60% weight and social factors a 15% weight, or vice versa. MSCI tends to overweight governance because it’s easier to quantify (board meetings, independent directors, etc.). Sustainalytics leans toward environmental because the data is more readily available.
I recall a conversation with a data vendor who joked that if you feed the same raw data into three ESG models, you’ll get three different portfolios. That’s because the models are built for different users. A long-term pension fund cares about transition risk. A hedge fund wants short-term controversy detection. A retail investor might just want a simple “green” badge.
One size fits all is a myth; one size fits *most* is already a stretch.
For a real-world illustration, consider the oil and gas sector. Shell received a top-tier ESG rating from one major agency in 2022, primarily because of its governance structure and carbon capture investments. Yet the same year, it was flagged by another agency as a “high risk” investment due to its fossil fuel revenue exposure. Same company, same year, completely opposite conclusions. If you’re an institutional investor trying to align with Paris Agreement goals, these ratings are not just unhelpful—they’re misleading.
## Data Quality: Garbage In, Gospel Out
Here’s a dirty secret: most ESG data is self-reported, unaudited, and inconsistent. Companies fill out questionnaires voluntarily. Some report everything; some cherry-pick. Some hire consultants to make their sustainability reports look pristine; others just copy-paste from last year with minor tweaks. And the rating agencies? They do their best, but they’re far from omnipotent.
In my day-to-day work, I’ve seen data gaps that would make an auditor weep. One of our portfolio companies in Southeast Asia discovered that its waste management data for two overseas factories was completely missing—not because they were hiding anything, but because no one had ever thought to collect it. The rating agency, without the data, assumed the worst and deducted points. The company appealed, but the process took months.
Data quality issues lead to false precision. An ESG score of 73.4 looks scientific, but if that score is based on 60% incomplete data, it’s closer to noise than signal. Some agencies now use AI and satellite imagery to fill gaps—forest cover, thermal emissions, shipping patterns—but even that is imperfect. A satellite can’t tell you if a factory is following labor laws. A natural language processing algorithm can’t detect whether that “strong anti-corruption policy” document is actually enforced.
The integration challenge here is twofold. First, we need common data standards (the ISSB framework is a promising step). Second, we need to accept that *missing data is itself information*. A company that doesn’t report social metrics is probably hiding something. But how do we penalize that consistently without punishing small firms that lack resources? It’s a fine line, and most agencies haven’t figured it out.
## Temporal Blind Spots: Past, Present, and Future Performance
Another dimension of divergence is *time*. Is an ESG rating a backward-looking score of what a company has done, or a forward-looking predictor of what it will do? Most agencies mix both, but with different emphasis. Some give high marks for historical achievements—like having already cut emissions by 20%. Others focus on future commitments—like a net-zero pledge for 2040. The problem? The future never arrives on time.
I remember a client in the aviation sector. They had a stellar “transition readiness” score because of their sustainable aviation fuel (SAF) purchase agreements. But those agreements were conditional—they only kicked in if the fuel passed certain lifecycle assessments. When those assessments were delayed by two years, the company’s actual emissions trajectory stayed flat. The rating agency didn’t update its score for another fiscal year. Meanwhile, another agency had already downgraded them. The market saw two conflicting signals for months.
Ratings are snapshots, but investors need moving pictures. This temporal mismatch is particularly severe in social metrics. Workforce diversity data, for example, reflects the past decade of hiring practices, not today’s initiatives. A company could have a terrible 2020, change its entire HR policy in 2021, and still show low diversity scores in 2023. Agencies that weight recent trends more heavily will diverge sharply from those that use trailing averages.
The integration solution isn’t to abandon historical data—that would be reckless. But we need a standard for *dynamic scoring*. Perhaps a baseline score plus a “momentum factor” that captures recent improvements. Some agencies (including a few we work with) are experimenting with this. It’s early, but the direction is promising.
## Regulatory Patchwork: When Governments Disagree, Ratings Follow
Now, let’s add regulators into the mess. The European Union has CSRD (Corporate Sustainability Reporting Directive), which is heavy on double materiality—you must report both how sustainability affects your business *and* how your business affects the world. The US SEC has proposed climate disclosure rules, but they’re lighter and mostly focused on financial materiality. China has its own disclosure frameworks that emphasize regulatory compliance over voluntary ESG metrics.
These laws shape what data is available, which in turn shapes what ratings can be calculated. A company listed in both London and New York faces conflicting reporting requirements. Some agencies base their scores on the *maximum* data available; others use the *minimum* required by the company’s home market. The result is a mess.
We are heading toward regulatory convergence, but slowly. The ISSB’s IFRS S1 and S2 standards are a huge step forward. But even those are voluntary in many jurisdictions. Until regulators align, rating agencies will keep pulling from different data pools. And that means the same company will get different scores simply because of where it’s headquartered.
I’ve seen this at BRAIN TECHNOLOGY LIMITED when we advise clients on cross-border listings. A European firm with a full CSRD report might suddenly see its rating *drop* when it lists in the US, not because it got worse, but because the US-focused agency applies less stringent baselines for comparison. This is a perverse outcome. The company didn’t change. The yardstick did.
## The Integration Imperative: From Parallel Systems to Layered Consensus
So where does this leave us? Throwing our hands up and declaring ESG ratings useless would be a mistake—they’re still the best proxy we have, and *any* information is better than none. But the path forward isn’t a single super-rating. It’s a layered, multi-signal approach.
Think of it like weather forecasting. You don’t consult just one app. You look at the radar, the satellite, the local station, and your own eyes. Each has strengths and biases. The skill lies in synthesizing them. Similarly, for ESG, we need to stop treating ratings as verdicts and start treating them as *inputs*.
One practical method is **composite scoring with transparency**. Take five or six major ratings, normalize them, and weight them based on each agency’s methodology strengths. If MSCI is better on governance and Sustainalytics is better on social controversies, why not build a composite that reflects that? At our firm, we’ve been building exactly this kind of internal modelfor clients. It’s not perfect, but it eliminates wild outliers and reduces the “single agency blind spot” problem.
Another integration path is **proprietary vs. public data fusion**. We increasingly pull in alternative data—news sentiment, satellite imagery, employee reviews on Glassdoor—and merge it with official ESG ratings. This gives us a richer picture. A company might have a high governance score from MSCI, but if employee reviews indicate sweeping staff turnover, we know there’s a social problem lurking. The integration isn’t mathematically elegant, but it’s practically effective.
The future might also hold *dynamic rating hubs*—centralized platforms where companies upload their raw data once, and multiple agencies apply their own models on top. This would solve the data gap problem and make comparisons more apples-to-apples. To be honest, the technical side is feasible. The politics are the bottleneck.
## Case Study: The Hydrogen Startup That Confused Everyone
Let me share a concrete example from our daily work. We had a hydrogen fuel cell startup in 2023. They approached us because their ESG scores were all over the place. One agency gave them 88/100 for environmental impact—hydrogen is clean, right? Another gave them 41/100, citing high energy intensity in their production method (they used grey hydrogen, produced from natural gas). A third agency ignored environmental entirely and scored them low on *governance* because their board lacked independent directors.
We ran a deep-dive. The truth? They were a classic *transition company*. Their current operations were not clean, but their technology roadmap was genuinely transformative. Existing rating methodologies couldn’t handle this nuance. One agency was too backward-looking; another was too future-oriented. Neither could integrate the company’s *trajectory*.
We built a custom “transition score” for this client, blending current performance with a credible transition plan. That score—78/100—helped them secure a green bond issuance. But here’s the irony: their official ESG ratings from the big three remained low and divergent for another year. It took a bespoke model to reveal what all three agencies were missing. This isn’t a failure of any single agency. It’s a failure of the system to accommodate non-linear change.
This case reinforced my belief that **integration isn’t about making ratings identical; it’s about making complementary ratings visible simultaneously.** Investors need a dashboard, not a single needle.
## Conclusion: Toward a Pragmatic Pluralism
We’re not going to wake up tomorrow with a universally accepted ESG rating. And honestly, maybe we shouldn’t. The divergence we see today is partly the result of healthy debate about what “good” looks like across industries and cultures. The problem is not the existence of differences; it’s the *opacity* of those differences and the false authority we assign to any single number.
My recommendation is threefold. First, regulators should push for mandatory, baseline data disclosure so that at least the *inputs* are consistent. Second, rating agencies should publish their methodology weights with more granularity—clients deserve to know why a score is what it is. Third, investors should adopt multi-rating frameworks, treating any single score as one signal among many.
At
BRAIN TECHNOLOGY LIMITED, we’ve already started building this into our products. We call it “ESG Signal Fusion.” It’s not flashy, but it delivers better risk assessments for our clients. Because at the end of the day, a rating is just a tool. The real goal is better decisions. And better decisions come from embracing complexity, not hiding from it.
As we move toward 2025 and beyond, I expect to see more regulators adopting ISSB standards, more agencies using AI for alternative data, and more investors demanding transparency over simplicity. The convergence will be messy. But it will be progress. And in this industry, that’s all we can ask for.
---
## BRAIN TECHNOLOGY LIMITED’s Perspective
At BRAIN TECHNOLOGY LIMITED, we’ve spent countless hours wrestling with the differences and integration of ESG ratings, not just as a theoretical exercise but as a core operational challenge. Our work in
financial data strategy and
AI finance has taught us that
the true value of ESG integration doesn’t lie in picking a “winner” rating agency, but in building adaptive systems that can process divergent signals simultaneously. We’ve learned that the ‘noise’ of conflicting ratings often carries more informational value than the consensus score, because it highlights areas of methodological disagreement—areas where the company’s strategy is either ahead of or behind the curve. Our AI-driven models now treat rating divergence itself as a feature, not a bug, using it to flag potential controversies or pioneering practices that single models miss. We believe the future belongs to platforms that synthesize, not dictate. For us, and for our clients, the goal isn’t to end the debate of ESG ratings—it’s to make the debate more productive.