Skip to content

WordPress

PHP serialized data repair

Repair serialized data that a database search-replace broke, and see what is actually inside it.

Everything runs in this browser tab. Nothing you paste is uploaded, logged, or stored — reload the page and it is gone.

Your widgets have vanished. The theme customisations reset themselves. A plugin’s settings page is empty even though the row is still sitting there in wp_options. And you ran a find-and-replace on the database about ten minutes ago.

Nothing is lost. Here’s what happened, and how to get it back.

What actually broke

PHP’s serialize() writes the byte length of every string in front of it:

s:5:"hello";

That number is the only thing describing where the string ends. unserialize() reads the 5, takes five bytes, and expects a closing "; right after. When it finds something else it gives up and returns false for the entire value, not just the one string that’s wrong.

A SQL REPLACE() changes the bytes and leaves the number exactly as it was:

UPDATE wp_options SET option_value = REPLACE(option_value, 'http://old.test', 'https://new.example');

http://old.test is 15 bytes and https://new.example is 19. Every string that contained the old URL is now understating its own length by four, and every one of them has become unreadable. WordPress gets false back, falls through to its default, and your settings look deleted. They aren’t. They’re intact and mislabelled, which is a much better problem to have.

Why the obvious fix makes it worse

The regex everyone reaches for first looks something like this:

preg_replace_callback('/s:(\d+):"(.*?)"/', $fix, $data);

Two things are wrong with it.

The lengths are byte counts, not character counts. café is four characters and five bytes. An em dash is one character and three bytes. An emoji is four. So a repair that counts characters introduces fresh corruption into every payload containing non-ASCII text, which on a real site is most of them. This tool counts UTF-8 bytes the same way PHP does.

The second problem is that string content can itself contain ";. A non-greedy match stops at the first one it sees, somewhere in the middle of your data, and shreds the structure below it. So this is a proper recursive-descent parser instead. It walks the actual tree, and when a declared length doesn’t land on a terminator it scans ahead for candidates, throws away any that aren’t followed by a legal token, and picks whichever one has a true byte length closest to what was declared. A search-replace shifts a length by a little rather than by an arbitrary amount, so “closest” turns out to be a reliable guess.

Nothing here gets uploaded

The parser runs in this page. No request, no endpoint, nothing logged.

That matters more for this tool than for most, because what you’re about to paste came out of a production database and could easily contain API keys, licence codes, SMTP credentials or customer records. Every other repair tool I looked at wants you to POST that to their server. Open the network tab and check this one before you trust it. Then close the tab and it’s gone.

Getting the fix back into the database

Copy the repaired output and write it back with a parameterised query rather than string concatenation:

UPDATE wp_options SET option_value = %s WHERE option_name = %s;

If the row is a big one, confirm it round-trips before you commit anything:

var_dump( unserialize( $repaired ) !== false );

Not doing this again

Don’t run REPLACE() against a WordPress database. Use something that unserializes, substitutes and reserializes properly:

wp search-replace 'http://old.test' 'https://new.example' --all-tables --dry-run

Drop --dry-run once the report looks sensible. WP-CLI walks into serialized values and recalculates the lengths instead of leaving them stale. Run wp db export first either way, because this kind of corruption is easy enough to fix one row at a time and genuinely miserable across a whole table. If WP-CLI isn’t available, the Better Search Replace plugin does the same job from the admin.

Some things that come up

Can this recover data that was actually truncated? No. If the bytes are gone they’re gone. This fixes the case where the content survived and the length is wrong, which is what search-replace corruption looks like. A value cut off because the column was too narrow needs a backup.

My repaired string still fails. Check for a latin1 / utf8mb4 collation mismatch. If the table charset changed at some point then the bytes themselves were re-encoded, and the wrong lengths are a symptom rather than the cause. Sort the collation out first, then come back.

Nested serialized data. Fairly common in wp_options: a serialized array holding a string that is itself serialized. Repair the outer layer, copy the inner string out, repair that separately, paste it back. The depth counter tells you how many layers you’re dealing with.

Objects. O: records parse fine and keep their class names. Bear in mind that unserialize() on an object needs the class to be loaded, so something left behind by a deactivated plugin will still fail even when the syntax is perfect.

Published