I burned a night on a 'Connect bug' that was two systems writing the same field.
I burned a night on the wrong villain
I was sure Sitecore Connect was dropping updates. Logs looked noisy. Retries looked heroic. The title field still flapped. Morning came. Coffee was bad. The map was worse. Content Hub and XM both thought they owned title. Last write won. Then the next job wrote it back. That's not a connector bug. That's two cooks.
I told the room it was Connect because Connect is easy to blame. Blaming a vendor is a meeting skill. Mapping fields is work. I did the meeting skill first. Don't copy that.
Idempotency in plain terms
Same payload twice should not create a second item. If it does, your key is wrong. I used display name once. Display names change. People add a space. Someone 'fixes capitalization.' Now you have twins. Use a stable ID from Hub. If you don't have one, stop the recipe. Don't 'just ship' and clean up later. Later is two thumbnails that look identical and aren't.
Replay the message in a lower environment until you're bored. Bored is the correct feeling. If the item count goes up, your key is soft. Soft keys plus retries is how clones are born.
Field ownership, written down
Pick a source per field. Write it on a sheet. Tape it. Hub owns DAM URLs. XM owns presentation. If both write body, you'll chase ghosts for a week and call it 'eventual consistency.' It isn't. It's a fight.
The sheet is not glamorous. Neither is a 2 a.m. flap. I keep the sheet next to the recipe name because I will forget. I forget names. I don't forget being woken up.
- Stable external ID.
- Upsert, not blind create.
- One owner per field.
- Replay twice before prod.
Retries
Retries are good. Blind create on retry is not. If Connect can't upsert, dead-letter it. I like retries. I like unique items more. Timeouts will happen. Timeouts plus create is how listing pages grow extra images.
We retried a failed Hub job. Two assets. Same title. Different IDs. The page used the ugly one. Of course it did. Murphy doesn't need a license.
What to check before you rewrite the recipe
Ownership map. Key. Replay count. Then the recipe. I rewrote the recipe first last time. I made a prettier wrong recipe. Pretty wrong still clones.
Ask one question in the standup: who owns this field. If two hands go up, stop. Don't add a transform. Transforms on a fight make a louder fight.
A night I don't want again
Phone on the nightstand. Alert on a publish. Title oscillating. I logged in from bed. That's a choice. The better choice was the sheet we didn't have yet. We have it now. It's stained. Good.
If your team pages you for field flaps, you don't have an observability problem first. You have an ownership problem. Logs will show the flap. They won't pick the owner. You have to.
Further reading, the useful kind
Sitecore's own docs on Connect recipes and Hub identifiers. Read the identifier page twice. Then read your recipe. Then the sheet. In that order. I used to skim identifiers. That's how display name got in.
If your lawyer cares about which system is source of truth, show them the sheet, not a architecture slide with clouds. Clouds don't testify. Dates on a sheet do.
Checklist I actually use
Replay the same message. Confirm no clone. Confirm the field you care about didn't flap. If it flapped, stop adding retries. Fix the map. Confirm upsert. Confirm the ID is not a name. Confirm a human who isn't me can find the sheet.
Ship only when replay is boring. If replay is exciting, you're not done. Exciting is for demos. Demos don't run at 2 a.m.
Monday
Print the field list. Put initials in the owner column. If a cell is blank, that field doesn't sync this week. Blank is allowed. Guessing is not.
Then turn off one retry path that creates. Watch the dead-letter. Dead-letter is not failure. Silent clone is failure. I'd rather an item wait than a twin go live.
If You Only Do One Thing
Turn the Hub job to manual for a week. Watch who yells. If nobody yells, you were syncing for a dashboard. If someone yells, you found the real consumer. Keep that consumer. Kill the rest of the schedule until it earns a slot.
I still have the Teams ping from the first silent Monday. It said 'is Hub down.' Hub was up. The job was off. That's the ping I wanted.
What I Won't Do
I won't add a second recipe until the first one has a replay runbook with a name on it. Second recipes are how we got the Tuesday pile.
The Tuesday Pile, Again
I opened Hub and counted rows that shouldn't exist. Twelve. Same promo, different timestamps. The recipe had no client key. Hope had been the key. Hope duplicated a homepage module. Marketing asked if the CMS was haunted. It wasn't haunted. It was scheduled.
We stopped the schedule. The pile stopped growing. That's the only graph I printed. It goes down when the job is off. Fancy monitoring said the job was healthy the whole time. Healthy and wrong is a genre.
Who Yelled
The catalog owner. Not the dashboard owner. Catalog people notice when a product shows twice. Dashboard people notice when a tile is green. Green lied. The catalog didn't.
I keep the catalog owner on the runbook now. If they aren't in the runbook, the job doesn't run. That's rude. Rude is cheaper than twelve rows.
Replay Day
We ran a replay on purpose after the key went in. One row. I stared at it like it might split. It didn't. I still don't trust it. I trust the count after the replay. Counts after replays are the religion.
Further reading
- Sitecore Connect Overview
- Content Hub Synchronization
- Idempotency in APIs
- Managing Field Ownership
- Retry Logic Best Practices
Actionable checklist
- Define clear field ownership for all synced content.
- Perform idempotency checks in your API calls.
- Set retry windows based on content importance.
- Regularly review and update ownership rules.
- Log sync actions to keep track.
- Test syncing processes in a staging area.
- Monitor sync success rates and any errors.
- Include a manual review for critical updates.
- Document your API endpoints and their functions.
- Regularly adjust integration strategies as needed.