A garment supplier evaluation scorecard turns opinion into evidence. It tracks on time delivery, first pass quality, measurement accuracy, sample turnaround, responsiveness, price stability, compliance status and flexibility on repeats, so buying decisions rest on a pattern of performance rather than on the last order anyone remembers.
Most brands manage suppliers by memory. A factory that missed a date in March is quietly dropped, while one that produced a difficult style well is never credited because nobody recorded it. The result is a supply base shaped by recent events and personal relationships rather than by how each factory actually performs.
A scorecard does not need software to be useful. What it needs is a short list of measures that are collected the same way every time, and a rule for deciding when a bad result means a bad supplier.
Why Move From One Off Buying to a Managed Supply Base?
One off buying works while volumes are small. Each season you find a factory, negotiate a price, run the order and move on. Once you are placing repeat business across several categories, that approach starts to cost money in ways that do not appear on an invoice.
- Every new supplier restarts the learning curve on fit, fabric and packing.
- Prices cannot be compared meaningfully because the specification changes each time.
- You have no basis for asking for better terms, since you cannot show what you have placed.
- Compliance and documentation have to be rebuilt from scratch for each factory.
- You carry the same risk on every order, because nothing accumulates.
A managed supply base is the opposite. A smaller set of factories, each measured over time, each with a known strength. That does not mean loyalty for its own sake. It means knowing which factory is right for which product and having evidence for the choice.
What Should a Supplier Scorecard Measure?
Keep the list short enough that it is actually completed after every order. Eight measures cover apparel supply well.
- On time delivery, measured against the date agreed at order confirmation, not the date the calendar was revised to.
- First pass quality rate, meaning the proportion of final inspections passed without rework or re inspection.
- Measurement accuracy, the proportion of checked points falling inside the agreed tolerance across graded sizes.
- Sample turnaround, from receipt of a complete comment to dispatch of the corrected sample.
- Responsiveness, a simple judgement of how quickly and completely questions are answered.
- Price stability, whether quoted prices hold and whether changes are explained by fabric or currency movement.
- Compliance status, whether audits, certificates and documentation are current without chasing.
- Flexibility on repeats, whether the factory takes smaller repeat quantities and short lead time top ups.
Two rules make the scores comparable. Define each measure in one sentence and never change the definition mid year. And weight the measures according to your business, because a fast fashion buyer and a premium outerwear buyer do not value the same things.
How Do You Score On Time Delivery Fairly?
On time delivery is the measure most often recorded incorrectly, because the reference date moves. A factory that agreed a date, then received fit approval three weeks late, then shipped two weeks late has not necessarily failed. Scoring it as a miss without recording the cause produces a scorecard that punishes suppliers for the buyer's own delays.
Record two things instead of one. Record whether the goods shipped on the originally confirmed date, and record who caused each change to the calendar. Over several orders this separates two different problems: a factory that plans badly, and a buying team that approves slowly.
It is also worth recording how the supplier communicated the delay. A factory that flags a fabric problem six weeks out is managing your order. A factory that reports the same problem three days before shipment is managing its own.
How Do You Measure Quality Across Orders?
Quality becomes comparable only when it is measured the same way each time. That means a defined inspection standard, a defined sample size and a defined defect classification, applied by the same team or to the same written rules across every supplier.
Two figures do most of the work. First pass quality rate tells you whether the factory delivers conforming goods without a second attempt. Measurement accuracy tells you whether the pattern and the cutting are under control, which is usually a better predictor of future problems than defect counts, because measurement drift is systemic rather than accidental.
Record defect types as well as defect counts. A factory with recurring stitching faults has a training or machinery issue. A factory with recurring fabric faults may have a mill problem that is not its fault at all. The same score can point at completely different actions.
How Do You Tell a One Off Failure From a Systemic Problem?
A single bad order does not mean a bad supplier. Fabric arrives faulty, a dyehouse misses a lot, a machine fails, a key person leaves. Dropping a capable factory after one incident destroys the very history the scorecard exists to build.
Three questions usually settle it. Has this type of failure happened before with this supplier, on a different style and a different fabric. Was the root cause inside the factory's control, or upstream at a mill or trim supplier. And how did the factory respond once the problem was known, in terms of speed, honesty and the cost it absorbed.
If the answer is that this is the first occurrence, the cause was upstream and the factory absorbed the correction, the score should record the incident without changing the relationship. If the same failure appears on unrelated styles, the problem is in how the factory works and no amount of goodwill will fix it.
How Do You Record This Without Heavy Systems?
A spreadsheet is enough for most brands, and a spreadsheet that is completed is better than a system that is not. One row per order, one column per measure, one column for the cause of any delay and one for free text notes. Suppliers appear on a second sheet as a rolling average over the last set of orders.
Three habits keep it alive. Complete the row within a week of shipment, while the detail is still known. Do not score a supplier on fewer than about three completed orders, because a single result is noise. And review the sheet at a fixed point in the buying calendar, before supplier allocation for the next season, rather than only when something has gone wrong.
Keep the scoring scale simple. A five point scale per measure is easier to apply consistently than a percentage, and consistency matters more than precision.
Should You Consolidate Volume or Dual Source?
Once you can see performance, consolidation becomes a rational decision rather than a leap. Moving more volume to strong performers gives you better attention, better terms and a shorter calendar, because the factory already knows your fit blocks, your packing and your standard.
The risk is concentration. If one factory holds your best selling style and it has a fire, a labour issue or a capacity clash with a larger customer, you have no route to market. Dual sourcing a key style means qualifying a second factory on the same style, with the same pattern and an approved pre production sample, and placing enough volume there to keep the relationship real.
A workable balance for many brands is to consolidate the bulk of a category with a lead supplier while keeping the styles that carry most of the revenue qualified in two places. Working across a network rather than a single factory makes that easier, because a second qualified source can be identified without starting a supplier search from nothing.
How Does an Office on the Ground Collect This Data?
Most scorecards fail because the data is collected by questionnaire. A supplier asked to report its own on time performance will report the revised date. A buyer scoring quality from a distance is scoring the inspection report rather than the factory.
Data collected continuously is different. A team that sits in Turkey and visits the floor sees line loading, the state of the machinery, whether the pre production sample was made on the line or in the sample room, and whether a promised delivery has any fabric behind it. Tekstil A.Ş. Global has worked in Turkish textiles since 1980, runs a 48 person team from Atasehir, Istanbul, and works across a network of 2,000+ verified member manufacturers, with an in house QC team recording quality and measurement data order by order.
The commercial model matters here too. Because the company is paid by the buyer and not by the factory, there is no incentive to record a supplier more favourably than its performance justifies, and no reason to steer volume towards a particular factory. That is the difference between a scorecard that guides your buying and one that confirms what someone else already decided.
Frequently Asked Questions
How many measures should a supplier scorecard have?
Enough to be meaningful and few enough to be completed after every order. Around eight works well for apparel: on time delivery, first pass quality, measurement accuracy, sample turnaround, responsiveness, price stability, compliance status and flexibility on repeats. Define each one in a single sentence and keep the definition fixed.
How many orders before a supplier score is meaningful?
Around three completed orders is a reasonable minimum. A single order reflects the style, the fabric and the calendar as much as the factory. Scoring after one result tends to reward suppliers who happened to receive an easy order and penalise those who took on a difficult one.
Should a supplier be dropped after one late delivery?
Not on its own. Check whether the same type of failure has happened before on a different style, whether the root cause was inside the factory's control or upstream at a mill, and how quickly and honestly the factory responded. Repeated failures on unrelated styles point to a systemic issue.
Do I need software to run a supplier scorecard?
No. A spreadsheet with one row per order and a rolling supplier average is sufficient for most brands. The discipline matters more than the tool: complete the row within a week of shipment and review the sheet before supplier allocation for the next season.
How do I score on time delivery when the buyer caused the delay?
Record two figures. Whether the goods shipped on the originally confirmed date, and who caused each change to the calendar. Over several orders this separates a factory that plans badly from a buying team that approves samples and lab dips slowly.
Is dual sourcing worth the extra cost?
For styles that carry a significant share of revenue, usually yes. Qualifying a second factory on the same pattern with its own approved pre production sample gives you a route to market if the lead supplier loses capacity. Place enough volume there to keep the second source genuinely ready.
Can a sourcing office collect scorecard data for us?
Yes, and it is more reliable than a questionnaire because the data comes from the factory floor rather than from the supplier's own reporting. Tekstil A.Ş. Global's in house QC and merchandising team records delivery, quality and measurement performance order by order across the supplier network.