When shared storage stumbles, everything leaning on it falls over at the same moment, and that is what makes these calls different from the rest of the week. SAN estates get reconstructed here — LUNs offline, pool metadata broken, VMFS datastores that will not mount — for companies and IT providers right across Norfolk, from one shelf to a room full of them. Enterprise work is core bench trade rather than an occasional favour.
Every san job is diagnosed free. The quote follows in writing, fixed, before a screwdriver is picked up.
No fix, no fee all jobs except electronic and mechanical failures, chip level work, DVR and Forensic jobs. Full pricing is on the data recovery cost page.
The first job on any san is matching the symptom to the fault — and after twenty-odd years, these thirty account for very nearly everything that comes through the door.
Enterprise shelves shed drives by the handful rather than one at a time. Every casualty gets individual repair and imaging before any group arithmetic resumes.
The block device your servers were relying on has gone. It gets rebuilt from the bottom up: RAID groups first, then pool metadata, strictly in that sequence, because doing it out of order proves nothing.
The map from blocks to volumes has failed. It gets reconstructed out of the raw on-disk structures it had been describing all along, which are still sitting there.
The management plane died mid-flash while the data sat perfectly still behind it. Work proceeds straight against the drives and ignores the controller entirely.
Pairs fail at the same time considerably more readily than the datasheet implied. Recovery goes drive-direct, which has the advantage of never having read the datasheet.
The virtualisation layer collapses first and takes everything visible with it. Datastores are parsed out and each virtual disk is then extracted and repaired on its own merits.
Seconds to do and days to regret. What survives depends on what has been written since, so freezing the platform improves the arithmetic immediately.
One broken link brings down every volume that depended on it. The chain is re-forged at image level, one link at a time.
Promised capacity exceeding real capacity corrupts every thin volume fed by the pool at the moment it runs out. Rebuilt from the images, with the accounts standing up to inspection afterwards.
Expired cache batteries plus a single outage leave parity contradicting data right across a group. At image level that contradiction can be resolved affordably rather than destructively.
NetApp aggregates and EMC pools lose track of where their volumes live. Reconstruction climbs up from the RAID groups until they remember.
520, 524 and 528-byte formats blunt ordinary tooling by design. The lab fabric here reads them natively rather than through improvised adapters.
HPE's small-business staple arrives with its vDisk structures torn. It is a reconstruction this workshop has rehearsed to the point of boredom.
V3700 and V5000 units drop extents after a power event. They are remapped from the drives, extent by patient extent, until the volume is whole.
The box lost the target definition and the volume behind it rarely dies alongside. Extracted and re-presented over plumbing that can be trusted.
Corrupt the index and every reference to every shared block fails simultaneously. Rebuilt until shared blocks resolve properly again.
Both controllers wrote independently through a botched handover, so there are now two versions of events. They are reconciled at image level, timestamped and audited.
Arrays past end of support fail with no official route left. The lab route does not expire, because data has never read a support contract.
Replacements running different firmware misbehave inside elderly groups in ways that look like drive failures. Stabilise, image, then reconstruct, in that order.
Masking or fabric edits blind all the hosts at once while the volumes underneath sit in perfect health. The mapping gets rebuilt and handed back to its owners.
Same-lot drives meeting an hours-counter firmware defect fail inside days of each other, which is a known industry embarrassment rather than bad luck. The survivors get imaged at speed before they join in.
Nearline SATA sitting behind SAS interposer boards fails at the adapter far more often than at the drive itself. The adapters are bypassed and the drives imaged bare.
When the fast tier dies, every tiered volume loses the blocks it was using most, and the cold tiers on their own make no sense. The cold tiers plus the tiering metadata rebuild whatever the quick layer was holding.
A zoning slip without a clustered file system produces interleaved corruption from both directions. The two write histories are separated from images, operation by operation.
Service processor dead, credentials retired along with a contractor. The array will not converse and the drives have no such reservations, so extraction proceeds without a password.
Space set aside for snapshots is finite, and most platforms deal with running out by deleting the oldest snapshots automatically. The restore point somebody was counting on is gone before anyone is told it existed. What is left is carved from unallocated space on the imaged members.
A failover test, a mislabelled site, or a relationship resumed in the wrong direction, and the healthy copy is overwritten by the stale one at wire speed. The overwritten side is imaged and the older structures are recovered from underneath, which works far better the sooner it is stopped.
Arrays that encrypt at rest keep their keys in the controller or on a key manager. Lose both and the drives are readable and meaningless. That gets established at the diagnostic stage and said plainly, because there is no honest way around it.
Shelves are daisy-chained, and a failed cable, a failed expander or a shelf powered up in the wrong order can make a whole tier of disks disappear at once. It looks like catastrophic loss and is frequently a cabling fault. Establishing which comes before anybody starts rebuilding anything.
A new server is presented with several volumes and a well-meaning administrator initialises the one that looked spare. A quick format writes very little, so the volume underneath is largely intact — provided the host is disconnected before it starts filling the space it thinks it owns.
A SAN is a stack of abstractions sitting on one another: disks grouped into RAID sets, sets gathered into pools, and the pools cut into LUNs, on top of which sit the file systems and hypervisor datastores everybody actually uses. Recovery goes down that stack and then climbs back up in the same order. Copy every disk. Stand the RAID sets up again. Repair the pool metadata. Lift the LUNs out. Only then open the virtual machines and databases living inside them. The hardware is familiar enough — Dell EMC, NetApp, HPE MSA and EVA, IBM, Fujitsu — and NDAs, change control and whatever paperwork your organisation runs on slot into that sequence without slowing it down.
Nobody actually wants raw blocks handed back to them. What they want is the handful of systems the business genuinely runs on, working again by morning. Datastores are therefore brought up on the reconstructed LUNs, the VMDK and VHDX files pulled out and repaired, and the named SQL or Exchange workloads moved to the front so operations can restart while everything else copies quietly in the background. It is triage rather than heroics, done at a measured pace, because hurry in an enterprise recovery is how the second mistake gets made.
SAN recovery needs enterprise plumbing as well as recovery skill, and the workshop stocks both in depth:
Shelves and drives connect over our own fibre channel, SAS and iSCSI infrastructure, including the unusual sector sizes enterprise stock arrives with.
Whole RAID groups captured alongside one another, after any casualty drive has been through the mechanical and firmware benches first.
Vendor metadata rebuilt in the order it was made: groups, then pools, then LUNs, and then whatever file systems are riding on top.
Restored LUNs present their datastores again and each virtual disk is lifted clear and repaired on its own terms.
SQL Server and Exchange workloads named at the outset are verified early and moved up the queue, so the services that matter come back first rather than last.
Every original, drive or shelf, stays read-only from the day it arrives until the day it goes home. There is no stage at which that is relaxed.
Unusual badges rarely mean unusual hardware, because enterprise vendors rebadge from a shared parts bin and this workshop has met most of that bin more than once. Pools that have lost their maps are rebuilt from the RAID groups up; a pair of dead controllers simply routes the work drive-direct; eccentric sector sizes are read natively on the lab fabric rather than argued with. Send the labelled drives and let the shelf keep its rack bolts.
Before anything comes out of the rack, tag every drive with its shelf and its slot, because enterprise layouts punish guesswork about position more than most. Photograph the setup as found, then send only the tagged drives — shelves and controllers stay on site. Post them tracked and insured or use your own courier; there is no collection service here, and drives can be handed in at reception at our Cambridge location if that suits better.
Most of what reaches this bench arrived by tracked, insured post. It is the steadiest way to move a poorly drive, and a parcel posted in Norfolk is usually on the bench the next working day.
Is the drive still bolted inside a laptop, desktop, MacBook, iMac, server or CCTV / DVR recorder? The hard drive or SSD needs to come out first, and only the bare drive travels — taking drives out of machines is not something we do here. Storage soldered to a motherboard (Apple Silicon Macs, one or two very thin laptops) is the single thing beyond us: if it will not come out, it cannot come in.
↓ Print the shipping & booking-in form (PDF)
Mark the parcel for the attention of Cambridge Data Recovery. From Norwich it is about an hour and twenty down the A11, then two minutes off the A14 at Junction 32 — or next working day by tracked post. You hear from us as soon as it is booked onto the bench.
Unsure what to put in the box? Ring 0800 689 0668 before you seal it, or run the free online diagnostic.
Diagnosis free, one figure written down, most work under no fix no fee. Start online, or ring us.