SitecoreAI: When Publishing Gets Stuck
114 Queued Jobs, One Failed Republish, and a Rather Unpleasant Afternoon

#sitecore #sitecoreAI #xmcloud #publishing #dailywork
Sometimes, things go wrong in a way you don't immediately notice.
Everything seems to be working. Editors continue updating content, publishing jobs appear in the queue, and nobody gets an obvious error message.
Except that nothing is actually being published.
That's pretty much what happened to us recently in a SitecoreAI production environment.
A large republish operation had been running for more than two hours before eventually failing. And somewhere along the way, the publishing queue had stopped processing new jobs.
By the time we started investigating, 114 publishing jobs were waiting in the queue.
Not exactly what you want to discover in a production environment with multiple sites and editors actively working ;-)
So, what happened?
It started with a republish operation for one of our sites.
Nothing particularly unusual about that. But this time, the job kept running. And running.
After approximately two hours and 13 minutes, it eventually ended with a Failed status.
Meanwhile, other publishing requests had continued to pile up.
Looking at the Publishing Jobs overview, we noticed something interesting:
The failed republish had processed thousands of items before failing.
All subsequent publishing jobs remained in
Queued.No other jobs were running.
Newly triggered publishing requests weren't being processed either.
And this wasn't limited to the site where the republish had failed. We were dealing with a multisite environment, and publishing was effectively blocked across the entire environment.
At that point, it was pretty clear that waiting a little longer probably wasn't going to solve the problem.
Just restart the environment?
That was obviously one of the first ideas. ("Reboot tut gut" gilt halt immer noch)
But there was a catch.
From previous experience, we knew that restarting the SitecoreAI CMS environment cwould clear pending publishing jobs.
And while a restart might get publishing working again, losing 114 publishing requests wasn't really an option.
Those jobs represented actual changes made by editors across multiple websites.
Sure, we could ask everyone to publish their content again.
But with that many jobs, across that many sites, created by different editors?
Not exactly my favorite recovery strategy.
So before touching anything, we needed a way to preserve the publishing requests.
A closer look at the Publishing API
Luckily, SitecoreAI provides a Publishing API.
And it turned out to be exactly what we needed.
The API allows you to retrieve publishing jobs, including their status and publishing options.
So instead of relying on the Publishing Jobs overview alone, we could get the information programmatically.
Using Postman, we retrieved the queued publishing jobs and saved the results as JSON.
One little detail: the API responses are paginated.
Which means that retrieving the first page and saving the JSON isn't necessarily enough. You need to make sure you've actually collected all the jobs.
In our case, we ended up with 114.
More importantly, the responses contained the information we needed to recreate the publishing requests later: which items were targeted, which languages were selected, and whether the original jobs included related items or subitems.
Well, technically, not everything was quite as straightforward as I initially expected. There are a few details around job IDs, item IDs, and request structures worth understanding.
But that's a topic for a separate technical article. ;-)
A backup is only useful if you can restore it
Having 114 publishing jobs saved as JSON felt reassuring.
But I wasn't entirely comfortable relying on that alone.
Could we actually recreate those jobs through the API?
So we picked a small, uncritical publishing request and tried it.
Using the Publishing API, we created a new job with the same item, language, and publishing options.
And it worked.
201 Created.
The new job appeared in the Publishing Jobs overview, with the expected settings.
Of course, it immediately joined the other jobs waiting in Queued.
But that was fine. At this point, we weren't trying to fix publishing yet. We just wanted to know whether our recovery approach would work.
We also tested canceling that newly created job through the API.
That worked too.
So now we had a reasonable plan:
Back up the pending publishing requests, restore publishing functionality, and recreate the jobs afterward.
Time for a restart
Before proceeding, we contacted Sitecore Support to confirm our approach.
The guidance was pretty clear: if a publishing job is blocking the queue, canceling it may help. If publishing remains stuck and no jobs are being processed, restarting the CM environment is an appropriate recovery action.
There was also an important confirmation:
Restarting the environment clears queued publishing jobs. They won't automatically resume afterward.
Good thing we'd saved them.
So we restarted the production CM environment through the Deploy App.
And sure enough, after the restart, the previously queued jobs were marked as Canceled.
All of them.
At least that part behaved exactly as expected.
But the real question was still unanswered.
Was publishing actually working again?
The moment of truth
We created another small test job using the Publishing API.
And this time, something different happened.
Instead of sitting in Queued, the job moved to Running.
Finally! ;-)
The publishing process was accepting and starting jobs again.
That was the confirmation we needed to proceed with the next stage of recovery: recreating the backed-up publishing requests and monitoring their completion.
Of course, there's still an important difference between successfully creating a publishing job and having it finish successfully. A 201 Created response doesn't tell you whether the actual publishing operation worked.
And after a failed republish, it's also worth checking whether an additional republish is needed to ensure the published content is consistent.
But the immediate blocker was gone, and we had a way to recover the pending work.
What I took away from this
Looking back, the most useful part of this incident wasn't actually the restart.
It was having a way to understand and preserve what was sitting in the publishing queue before doing anything potentially disruptive.
A few things I'll definitely keep in mind:
Don't assume a queued job will eventually start. If nothing is running and new jobs keep piling up, it's worth investigating sooner rather than later.
Don't restart blindly. Understand what will happen to pending publishing requests first.
The Publishing API is more useful than you might think. It's not just about triggering publishing programmatically. It can also be a valuable troubleshooting and recovery tool.
Always test your recovery approach. Saving the jobs was one thing. Verifying that we could actually recreate them was another.
A running job isn't necessarily a successful job. Monitor the results and consider whether additional publishing is needed to restore consistency.
Wrap-up
What started as a failed republish turned into 114 queued publishing jobs and an environment that simply refused to publish anything.
Not particularly enjoyable.
But also a good reminder of why it's worth knowing the tools available beyond the SitecoreAI user interface.
The Publishing API gave us a way to preserve the pending work, test our recovery approach, and move forward without asking dozens of editors to repeat their publishing actions manually.
Sometimes, the most important step in fixing a problem is making sure you don't lose anything while fixing it.
And since the API turned out to be quite useful, I'll follow up with a more technical article showing the actual requests, pagination, and how to recreate and cancel publishing jobs using Postman.
Because hopefully, the next time someone ends up with a publishing queue full of stuck jobs, they'll have a slightly more relaxed afternoon than we did. ;-)




