It's 9:15 AM. You've been at Halcyon for three weeks. Your manager, James, drops a file in your inbox with one line: "Meridian renewal SOV — can you get this modelled by Thursday?"
You open the attachment. It's an Excel file with four sheets — Americas, Europe, Asia-Pacific, Middle East & Africa. 150 locations. 13 countries. You scroll through the first sheet. Some rows look complete. Others have blank cells where you expected numbers. One row says "TBC" under the TIV column. One Japanese record has a value of 850,000,000 with no currency label. A Lagos entry has no address, no construction class, and no coordinates — just the words "Apapa Port Area near Tin Can Island."
James hasn't told you what to do with any of this. He just wants results by Thursday.
The Schedule of Values (SOV) is the single most important document in any cat model run. It is a structured list of every insured location in a portfolio — its address, physical characteristics, and insured financial values. Everything the model produces flows directly from what this file contains.
The SOV does not originate in the insurance industry. It starts with the insured, a company's property management team, a real estate division, a corporate risk manager and gets compiled from whatever systems that company happens to use to track their assets. Those systems were built for accounting or facilities management, not for insurance or cat modelling. The data is organised the way the company finds useful, not the way a hazard model needs it.
By the time the SOV reaches your desk, it has passed through at least two other pairs of hands. The insured's broker extracts the data and formats it for the submission. Then it comes to you. Every translation step is an opportunity for fields to be left blank, values to be reformatted incorrectly, or addresses to be swapped for head-office mailing addresses. This is not incompetence, it is the structural reality of how exposure data moves through the market.
A cat model is not a machine that turns bad data into reliable loss estimates. It is a machine that amplifies what you give it. A portfolio with 40% of locations at county-level geocodes, will produce loss estimates that are wrong by an unknowable margin. And the model will produce those estimates without warning. Knowing what good data looks like before you run the model is the core skill of an exposure analyst.
The chain of custody for an SOV determines where its quality problems come from. Understanding this chain tells you exactly who to contact when you find a gap, and what kind of gap to expect.
A complete SOV for cat model input covers five distinct categories of information. Not every field in every category is equally important — but knowing what the full picture should look like is what allows you to assess what is missing and how much it matters.
There is a critical difference between fields a model needs to run and fields it needs to run accurately. The model will process an SOV with missing construction data — it will simply apply a default. That default is almost always conservative (loss-inflating) and may bear no relationship to what the building actually is. Understanding this table is the foundation of every data quality conversation you will have.
| Field | Status | If Missing |
|---|---|---|
| Location ID | REQUIRED | Every row must have a unique identifier |
| Country | REQUIRED | Model cannot select the correct hazard database |
| Address or coordinates | REQUIRED | Location placed at country centroid : hazard score is meaningless |
| Total Insured Value (TIV) | REQUIRED | No financial output possible : model cannot calculate a loss |
| Currency | REQUIRED (multi-country) | All TIVs treated as base currency : catastrophic error at scale |
| Occupancy class | IMPORTANT | Model applies most common regional default : may be wholly wrong |
| Construction class | IMPORTANT | Conservative default inflates losses : direction is predictable but magnitude is not |
| Year built | IMPROVES RESULTS | Cannot distinguish modern code-compliant from pre-code : critical for earthquake |
| Number of storeys | IMPROVES RESULTS | Default height assumption selected : affects resonance period for earthquake, flood depth ratio |
| Roof shape | IMPROVES RESULTS | Gable default applied : conservative for wind; can move hurricane losses 20–35% |
| First floor height | IMPROVES RESULTS | At-grade default maximises flood vulnerability : may overstate significantly for elevated buildings |
The first thing you do when an SOV lands in your inbox is assess it, not model it. Before you touch the model, you need to understand what you are working with. The following twelve rows are drawn directly from the Meridian Portfolio. Each row illustrates a data quality pattern you will encounter repeatedly in your career.
| Loc ID | Location | Country | Address | City | Postcode | TIV (local) | Currency | Construction | Occupancy | Year Built | Lat | Lon |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MGA-001 | Tampa Bay Office Complex | USA | 1847 Harbour View Blvd | Tampa, FL | 33602 | 12,500,000 | USD | Steel Frame | Office | 2004 | 27.944 | -82.459 |
| MGA-015 | London Canary Wharf Retail | GBR | Unit 4 Canada Square | London | E14 | 3,800,000 | GBP | Concrete | Retail | — | — | — |
| MGA-021 | Osaka Warehouse Facility | JPN | 2-10-70 Namba-Naka Naniwa-ku | Osaka | 556-0011 | 850,000,000 | — | RC Frame | Industrial/Warehouse | 1998 | 34.661 | 135.501 |
| MGA-036 | Istanbul Mixed Use Tower | TUR | Büyükdere Cad. No:127 | Istanbul | 34394 | 45,000,000 | USD | Mixed | Mixed Use | 2008 | 41.07 | 29.01 |
| MGA-063 | Lagos Logistics Hub | NGA | Apapa Port Area near Tin Can Island | Lagos | — | 2,100,000 | USD | — | Commercial | — | — | — |
| MGA-068 | Santiago Office Tower | CHL | Av. Apoquindo 3600 Las Condes | Santiago | 7550000 | 6,400,000 | USD | RC Shear Wall | Office | 2015 | -33.415 | -70.598 |
| MGA-074 | Bogotá Hotel & Conference | COL | Carrera 7 No. 32-16 | Bogotá | 110311 | TBC | — | Concrete Frame | Hotel | — | 4.615 | -74.068 |
| MGA-079 | Munich Residential Apartments | DEU | Maximilianstraße 18 | München | 80539 | 5,750,000 | EUR | Masonry | Residential Multi-Family | 1963 | 48.139 | 11.577 |
| MGA-089 | Bangkok Retail Complex | THA | Ratchadamri Rd Pathum Wan | Bangkok | 10330 | 280,000,000 | — | RC Frame | Retail/Commercial | 2011 | 13.743 | 100.540 |
| MGA-099 | Manila Office Tower | PHL | Ayala Ave Makati City | Manila | 1226 | 350,000,000 | PHP | RC Frame | Office | — | 14.557 | 121.017 |
| MGA-104 | Dubai Free Zone Warehouse | ARE | Plot 35-B Jebel Ali Free Zone | Dubai | — | 8,900,000 | USD | Steel Frame | Industrial/Warehouse | 2019 | 25.00 | 55.11 |
| MGA-138 | Mumbai Residential Tower | IND | Andheri West, Mumbai | Mumbai | 400053 | 320,000,000 | — | RCC | Residential High-Rise | 2017 | — | — |
| 12 of 150 Meridian Portfolio locations. Full CSV available for download at the bottom of this lesson. Colour coding: Red = must resolve before modelling · Amber = affects accuracy · Green = acceptable. | ||||||||||||
MGA-001 (Tampa Office) is the benchmark, this is what a clean record looks like. Full street address, precise coordinates, valid USD TIV, Steel Frame construction, Office occupancy, year built. Every required and important field is populated. This is the standard every other record should be measured against.
MGA-021 (Osaka Warehouse) has one critical error that makes it unmodelable: no currency. The TIV of 850,000,000 is meaningless without knowing whether it is JPY or USD. At 2024 exchange rates, JPY 850 million ≈ USD 5.7 million. If the model treats this as USD 850 million, the location appears as the most valuable single asset in the entire portfolio, which it is not. This one missing field makes the financial output for this location completely unreliable until resolved.
MGA-036 (Istanbul Mixed Use Tower) has three compounding issues. "Mixed" is not a construction class, it is a placeholder that forces the model to apply a default. The coordinates are given to only 2 decimal places (±1 km resolution) in a city that sits on one of the world's most dangerous active faults, where soil conditions vary dramatically at sub-kilometre scale. And 2008 construction in Turkey places this building in a critical code-era boundary, post-1999 Marmara earthquake reforms require investigation to confirm compliance. USD 45 million TIV at risk, and the three most important variables for earthquake modelling are all uncertain.
MGA-063 (Lagos Logistics Hub) is near-total data failure. "Apapa Port Area near Tin Can Island" is a neighbourhood description, not an address. No postcode exists because Nigeria does not have comprehensive postal coverage in this area. No construction class, no year, no coordinates. The TIV is the only usable field. This record will run through the model and will produce a number based entirely on model defaults applied to a Lagos city centroid geocode. That number will be essentially meaningless as a representation of this specific location's risk.
MGA-089 (Bangkok Retail Complex) repeats the Osaka currency problem TIV of 280,000,000 with no currency. THB 280 million ≈ USD 8 million, which is plausible for a Bangkok retail complex. USD 280 million is not. The magnitude of the figure alone is diagnostic — and this is a skill worth developing: knowing the approximate order of magnitude of asset values in different markets so that currency errors become immediately visible.
When the SOV arrives, your first job is to understand what you are working with, not to fix it. A structured ten-minute diagnostic tells you what the data quality issues are, which ones are critical, and who to contact. Running the model before completing this step means producing results whose limitations you do not understand.
How many locations? Does that match the submission documentation? Sum the TIV column, does it agree with the stated total? If the submission says USD 500M and the SOV sums to USD 50M, something is missing or the currency handling is broken.
Are the columns you need for model input present? Is there a currency column? Are TIV components separated (building / contents / BI) or combined? Are there columns with unclear purposes? Note anything unexpected before you start cleaning.
For each key field, what proportion of records have a valid value? A field that is 95% complete is a manageable gap. A field that is 30% complete is a structural problem that may require you to pause the process and request more data before modelling.
Does the country distribution match the submission? If the cover note says "UK commercial portfolio" and 15 records are in Germany and Poland, you need to understand why before you run the model. A quick country pivot takes 90 seconds.
Sort by the field with the highest blank rate and look at the 10 worst rows. Are the problems random (a few missed entries) or systematic (an entire column blank)? Random gaps are fixable. Systematic gaps require either additional data or an explicit statement to the underwriter about what the model does and does not reflect.
A file where every field is populated with plausible-looking values can be more dangerous than one with visible gaps. If 90% of locations share exactly the same year built, or every commercial property is listed as "concrete frame" regardless of country, the data has been batch-filled with assumptions. The model will run without complaint. The results will be based on the cedant's guesses about their own data rather than actual property information. Always look for implausible uniformity, it is the hardest data quality problem to spot and the easiest to miss.
Data quality problems do not stay in the SOV, they propagate directly into every number the model produces. The relationship is proportional and unforgiving: a 30% TIV understatement produces a 30% understatement in modelled losses. Conservative construction defaults inflate losses in a direction you can predict, but by a magnitude you cannot precisely know without the correct data.
Consider MGA-036 : the Istanbul mixed-use tower, under three data quality scenarios. The location, the hazard, and the USD 45 million TIV are identical in all three. Only the data quality changes.
The same building. The same earthquake hazard. A 2.2× difference in modelled annual expected loss driven entirely by data quality. Across a 150-location portfolio where many records are at Scenario B or C quality, the cumulative impact on AAL and the EP curve can be substantial and the direction is almost always toward overstatement, because conservative defaults inflate losses.
This has a direct pricing implication. A portfolio with poor data quality will appear more expensive in the model than the same physical risk with complete data. If the underwriter prices on the model output without understanding the data quality, they may over-price a risk and lose it to a competitor who has better data or, in the other direction, accept a poorly understood risk believing it is adequately priced when the dominant uncertainty is actually in the input data.
The type of account determines what SOV quality is realistic to expect and therefore how you should approach gaps. A Lloyd's direct placement from a large multinational has fundamentally different data quality dynamics than a reinsurance bordereaux from a regional mutual insurer.
Commercial property direct & facultative (D&F) placements like Meridian tend to have the most complete data because the broker works closely with a single client and can obtain property-specific information. Full addresses, construction types, and year built are all obtainable, even if they require chasing. The Meridian SOV's gaps are fixable with a focused data request.
Treaty reinsurance SOVs represent an entire cedant's portfolio, potentially tens of thousands of locations. Data quality varies enormously by cedant sophistication. A well-run specialty insurer will have detailed location-level data. A small regional mutual may submit aggregated postcode-level data with no COPE information. At treaty level, the aggregate exposure quality matters more than individual location detail but systematic gaps in construction or occupancy data across thousands of records can materially bias portfolio-level results.
Delegated authority bordereaux submitted periodically by MGAs and coverholders typically have the weakest data quality of any SOV type. They reflect the data capture practices of potentially dozens of different MGAs, each using their own system with their own field definitions. A property syndicate relying on 20 MGA binding authorities will likely receive 20 different approaches to construction coding, 20 different address formats, and 20 different approaches to TIV components. Harmonising this data before model input is a significant exercise in itself.
You have three options when an SOV arrives with material data gaps. First: request the missing data from the broker and delay the model run. Appropriate when the gaps are in fields that are critical for the dominant peril and the account is large enough to justify the timeline. Second: run the model with current data, apply and document conservative assumptions for the gaps, and present the results as directional only, with a clear statement of what changes once the data improves. This is appropriate for initial pricing indications on time-sensitive placements. Third: exclude records with critical gaps from the run entirely, and flag their exclusion explicitly. Never choose between these options without telling the underwriter which one you have taken and why.
The model produces a number. The number looks plausible. No one knows 23 currency fields are blank and 15 addresses are head-office locations. This is the most common and the most invisible error.
A company insures a Birmingham factory but the broker holds the London head-office address on file. The model geocodes the risk to the City of London. Wind zone, flood zone, and ground motion hazard are all wrong, silently.
"Concrete" is a material, not a structural system. Unreinforced concrete masonry and modern RC shear wall are both "concrete" with earthquake vulnerability factors 7–10× apart. Always investigate ambiguous construction descriptions for seismic-zone locations.
A 10-storey office building insured at USD 500,000 and a single-family home insured at USD 20 million are both wrong. A per-square-metre plausibility check against market benchmarks takes two minutes and catches order-of-magnitude errors before they enter the model.
Every gap filled with a default or assumption by you or by the model, must be in the run log. Six months later, when a loss occurs and the model output is scrutinised, you must be able to explain exactly what data was used and what was assumed.
On renewal business, always treat the new SOV as a fresh diagnostic exercise. Portfolio composition changes, broker teams change, and data quality can move in either direction between renewal years.
A structured list of all insured locations with their physical characteristics and insured values. The primary input document for every cat model run.
The maximum amount insurable at a location, typically the sum of building, contents, and business interruption values. The direct multiplier on every damage ratio the model produces.
Construction, Occupancy, Protection, Exposure : the four standard fields describing a building's physical risk characteristics. Primary drivers of vulnerability curve selection.
The periodic statement submitted by a coverholder or MGA listing all risks bound under a delegated authority. The functional equivalent of an SOV for binding authority business.
A value a cat model applies when a required field is missing. Defaults are calibrated to be conservative, they inflate modelled losses rather than understate them.
Any exposure field beyond basic COPE that refines vulnerability curve selection — year built, number of storeys, roof shape, basement presence, first floor height.
The proportion of records in an SOV that have a valid value for a given field. The primary quantitative metric for data quality assessment before a model run.
The spatial precision of a location coordinate, street-level, postcode-level, city-level, or country-level. Resolution determines the accuracy of the hazard score the model assigns. Covered in depth in Module 2.
Five questions · 4 of 5 correct (80%) to pass