THE EVALUATION MACHINE: WHY FRAMEWORK SCORING PRODUCES PREDICTABLE WINNERS
The Framework Reckoning · Article 5 of 12
Hard FM · Soft FM · Public Sector Procurement · CCS RM6378
Article 4 showed that framework call off awards concentrate heavily in a small number of suppliers. This article examines the machinery that produces that outcome: the evaluation model. How quality is scored at framework appointment. How price is weighted at call off. Why certain question types consistently favour certain types of supplier. And whether the scoring mechanism is measuring what it claims to measure.
15–30
Scored quality questions per framework
4 Years
Duration of the permanent filter exclusion
60–70%
Typical quality weighting at appointment
Two evaluations, two different tests
Framework procurement involves two distinct evaluations. The first happens at framework appointment: the framework body evaluates suppliers against quality, capability, and sometimes pricing criteria to determine who gets on the panel. The second happens at call off: the buying authority evaluates suppliers on the panel against the specific requirements of the contract being awarded. These two evaluations test different things. The framework appointment evaluation tests whether the supplier can describe a credible approach to FM delivery in general terms. The call off evaluation tests whether the supplier can deliver a specific contract at a competitive price. A supplier that excels at the first test, writing articulate, comprehensive quality responses, may not excel at the second, delivering an operationally competitive price for a specific estate. And a supplier that would be the best operator for a particular site may never get the chance to prove it if their framework appointment submission did not score highly enough. The structural problem is that the first evaluation acts as a permanent filter. A supplier that scores poorly on quality writing at framework level is excluded from every call off for the next four years, regardless of their operational capability. The framework evaluation does not test delivery. It tests description of delivery. Those are not the same thing.The framework evaluation does not test whether a supplier can deliver FM services. It tests whether a supplier can describe how they would deliver FM services. The gap between those two capabilities is where the evaluation machine produces its distortions.
The quality question problem
Framework quality evaluations typically consist of 15 to 30 scored questions. Each question asks the supplier to describe their approach to a specific aspect of FM delivery: mobilisation, workforce management, innovation, technology, sustainability, health and safety, quality assurance, continuous improvement, and social value. The question format is almost always the same: describe your approach to [topic], providing evidence of how you have delivered this on similar contracts. Suppliers respond with structured written submissions, typically 1,500 to 3,000 words per question, supported by case studies, process diagrams, and organisational charts.The Anatomy of a “Describe Your Approach” Evaluation
The format rewards three key criteria over operational delivery:
- Writing Quality: Articulate and polished corporate prose designed to hit precise evaluator matrices.
- Structural Clarity: Heavy reliance on massive, pre-written bid content libraries that allow swift adaptation.
- Volume of Case Studies: A deep pool of existing public sector contract profiles to easily format into evidence records.
What this format rewards
The describe your approach format rewards three things: writing quality, structural clarity, and volume of relevant case studies. A supplier with a dedicated bid team, a library of pre written content, and a portfolio of public sector FM contracts can produce a polished, high scoring response to any framework quality question. The content may be largely reusable across frameworks because the questions are largely the same across frameworks. A supplier without a dedicated bid team, without a content library, and with experience concentrated in a specific region or sector faces a structural disadvantage. Their operational delivery may be excellent. Their ability to articulate it in 3,000 words of scored prose, with case studies formatted to the evaluation panel’s expectations, may not be.The case study trap
Case study requirements compound the disadvantage. Framework evaluations typically require suppliers to provide named case studies demonstrating delivery of the service described. The evaluation panel scores the relevance, scale, and outcomes of the case studies provided. Suppliers with large portfolios of public sector FM contracts have a deep pool of case studies to draw from. Suppliers entering the public sector from the private sector, or growing from regional to national scale, may have strong delivery credentials but limited public sector case studies at the scale the evaluation expects. The case study requirement creates a feedback loop identical to the incumbency advantage described in Article 4. Suppliers who already hold framework contracts can cite those contracts as case studies in the next framework evaluation. Suppliers who have never held a framework contract cannot. The evaluation rewards past framework success, which reinforces the concentration pattern: the same suppliers score well because they have the case studies, and they have the case studies because they scored well before.
The Carbon Reduction Plan barrier
Since PPN 06/21, suppliers bidding for major government contracts must submit a qualifying Carbon Reduction Plan. This is a pass/fail requirement at framework appointment. A supplier without a compliant CRP is excluded before the quality evaluation begins. For national providers with dedicated ESG departments, producing a CRP is a routine compliance exercise absorbed into existing corporate infrastructure. For mid market and regional suppliers without ESG teams, producing a CRP to the required standard involves external consultancy, carbon footprint measurement, and target setting that represent a significant cost and capability barrier. The CRP requirement does not measure whether a supplier delivers low carbon FM services. It measures whether a supplier has the corporate infrastructure to produce a compliant document. Like the case study trap and the turnover threshold, the CRP requirement filters for organisational scale rather than operational capability.
Cross-Series Analysis: The SFG20 Reckoning series identifies a parallel dynamic in maintenance specification. SFG20 task schedules are used as the pricing basis in framework call offs, but SFG20’s generic task durations do not reflect site specific conditions.
Quality weighting at framework versus call off
The weighting given to quality versus price differs significantly between framework appointment and call off. At framework appointment, quality is typically weighted heavily: 60 to 70% quality, 30 to 40% price. This reflects the framework body’s interest in appointing suppliers who can demonstrate broad capability. At call off, the weighting often reverses: 30 to 40% quality, 60 to 70% price. This reflects the buyer’s interest in getting the lowest cost for a specific requirement. This inversion creates a structural tension. The supplier that scores highest on quality at framework level is the supplier that writes best. The supplier that wins at call off level is the supplier that prices lowest. If these are the same supplier, the system works. If they are not, the framework has appointed suppliers on the basis of writing quality and the call off awards work on the basis of price, and the two filters produce different winners.
The Weighting Inversion Dilemma
Appointing panels on writing skill and subsequently executing actual service awards purely on rock-bottom call-off pricing fragments delivery continuity long before assets are touched.
The quality floor problem
Once suppliers are appointed to the framework, the quality score that earned their place is rarely revisited. At call off, quality is often reduced to a pass/fail threshold rather than a differentiating score. Suppliers must demonstrate that they meet the minimum quality standard for the specific requirement. Beyond that threshold, price determines the winner. This means the detailed quality evaluation at framework level, the 15 to 30 scored questions, the case studies, the innovation narratives, becomes irrelevant at call off. The quality score got the supplier on the panel. It does not help them win the work. The framework evaluation and the call off evaluation are testing different things and rewarding different capabilities. The supplier that invested most heavily in quality writing at framework level may find that investment delivers no competitive advantage at call off, where price is the determining factor.Consensus scoring and moderation
Framework evaluations are typically scored by panels of evaluators using a consensus or moderated scoring methodology. Individual evaluators score each response independently. The panel then meets to agree a consensus score through discussion and moderation. Consensus scoring is designed to reduce individual bias and produce fair, consistent outcomes. In practice, it introduces a different dynamic: regression to the mean. Evaluators who score a response highly are moderated downward. Evaluators who score low are moderated upward. The consensus score gravitates toward the middle of the range. This compresses the scoring distribution and reduces the gap between the best and worst submissions. For established suppliers with experienced bid teams, compression is manageable. Their submissions are consistently good enough to score in the upper range after moderation. For new entrants or specialist suppliers whose submissions may be operationally strong but stylistically unconventional, moderation can be fatal. A single evaluator who scores them highly may be overruled by the consensus process. The evaluation rewards conformity to expected response structures as much as it rewards content quality.Consensus scoring is designed to reduce bias. In practice, it compresses the scoring distribution and rewards conformity to expected response structures. Suppliers who write the way evaluators expect to read will consistently outscore suppliers who write differently, regardless of operational capability.
Social value scoring: the unmeasured commitment
Social value is typically scored at 10 to 20% of the quality weighting at framework appointment. Suppliers commit to apprenticeships, local employment targets, SME subcontracting percentages, carbon reduction targets, community engagement hours, and mental health support programmes. These commitments are scored. Higher commitments produce higher scores. The problem, which Article 7 of this series examines in full, is that social value commitments made at framework level are rarely tracked through to delivery at call off level. The framework body scores the promise. Nobody audits the outcome. A supplier that commits to 50 apprenticeships at framework level and delivers 5 at call off level faces no scoring penalty on the next framework evaluation because the delivery data is not connected to the scoring system. This creates a rational incentive to overcommit. Suppliers who promise the most score the highest. Suppliers who promise conservatively and deliver fully score lower. The evaluation rewards ambition over delivery. The social value score becomes a writing exercise, not a performance measure. The ERIC Reckoning series examines an identical accountability gap in NHS estate data. ERIC returns are self reported by NHS trusts with limited independent verification. The data that drives capital allocation decisions has quality limitations the market has not confronted. The same structural problem applies to social value scoring in frameworks: self reported commitments without independent verification produce data that rewards reporting effort, not operational outcomes.What the evaluation machine actually measures
The framework evaluation machine is not biased in a legal sense. The questions are published. The scoring criteria are defined. The moderation process is documented. The evaluation panels are staffed by experienced procurement professionals who apply the criteria as written. But the machine is structurally predictable. It measures writing quality, not delivery quality. It rewards case study volume, not operational excellence. It compresses scoring distributions through consensus moderation, favouring conformity over capability. It weights quality heavily at appointment and price heavily at call off, creating a two stage filter that tests different things at each stage. And it scores social value commitments without tracking whether they are delivered. A supplier that understands the evaluation machine can optimise for it. Invest in a bid team. Build a case study library. Write to the expected structure. Commit ambitiously on social value. Price competitively at call off. This strategy will produce consistent framework appointments and consistent call off wins. Whether it also produces consistent operational delivery is a question the evaluation machine does not ask.
Structural Predictability
The evaluation machine is not biased. It is structurally predictable. A supplier that optimises for the evaluation will consistently outscore a supplier that optimises for delivery. That is the gap the framework model has not closed.
THE FRAMEWORK RECKONING · THE £120BN AUDIT OF UK FM PROCUREMENT · 12 ARTICLES
- IndexFull series index and analysis: baachurain.com/framework-reckoning
- PrevArticle 4: Who Actually Wins?
- NextArticle 6: The Pricing Illusion
- AdvisoryQuestions or support: hello@baachu.com