Skip to main content
Level 2
October 9, 2026
Solved

Sling jobs is getting stuck for 24+hours

  • October 9, 2026
  • 2 replies
  • 21 views

Our technology stack is : aem cloud / AEM Guides / java21

We have migrated from ams to aem cloud few weeks ago. After content publish, aem author generate pdf asynchronously for the published content. Sling jobs to generate pdfs are taking 24hr’s. AMS environment used to take 5-10minuntes to generate pdf’s.

 

Here is our  configuration:

{  "queue.name": "MyGPS Generate PDF Queue",  "queue.topics": [    "com/skylark/mygps/generatePDF"  ],  "queue.type":"ORDERED",  "queue.retries":1,  "queue.maxparallel":1.0}

*********

we have a custom servlet which prints all pending sling jobs and here is the snippet:

   {    "user": "michael.abc@abc.com",    "Payload": "/content/mygps/us/en/pi/public/trading/trading-basics/account-and-trade-restrictions/handling-fraud-related-restrictions-081821-jh_ditamap/Phonebooka17b4efc-716b-49cd-a909-480703ef2a74",    "topic": "com/abc/mygps/generatePDF",    "id": "2026/10/8/16/4/7564fa83-ecaf-4986-a1fb-da04e8eae38b_427",    "state": "ACTIVE",    "queueName": "MyGPS Generate PDF Queue",    "createdTime": "2026-10-08T16:04:41.257Z",    "pendingTimeSeconds": 79090,    "pendingTimeMinutes": 1318,    "pendingTimeHours": "21 hr 58 Minutes",    "retryCount": 0  },  {    "user": "michael.abc@abc.com",    "Payload": "/content/mygps/us/en/pi/public/support/compliance-and-risk/corporate-issues-policy/Compliance-policy-cybersecurity-121619-MR_ditamap/Phonebooka17b4efc-716b-49cd-a909-480703ef2a74",    "topic": "com/abc/mygps/generatePDF",    "id": "2026/10/8/16/4/7564fa83-ecaf-4986-a1fb-da04e8eae38b_428",    "state": "QUEUED",    "queueName": "MyGPS Generate PDF Queue",    "createdTime": "2026-10-08T16:04:41.648Z",    "pendingTimeSeconds": 79090,    "pendingTimeMinutes": 1318,    "pendingTimeHours": "21 hr 58 Minutes",    "retryCount": 0  },  {    "user": "michael.abc@abc.com",    "Payload": "/content/mygps/us/en/pi/help-desk/specialty/service-hd/advanced-service-hd/70fc6a38-46e2-4795-94b0-7023f3075e75_ditamap/Phonebooka17b4efc-716b-49cd-a909-480703ef2a74",    "topic": "com/abc/mygps/generatePDF",    "id": "2026/10/8/16/4/7564fa83-ecaf-4986-a1fb-da04e8eae38b_430",    "state": "QUEUED",    "queueName": "MyGPS Generate PDF Queue",    "createdTime": "2026-10-08T16:04:47.070Z",    "pendingTimeSeconds": 79084,    "pendingTimeMinutes": 1318,    "pendingTimeHours": "21 hr 58 Minutes",    "retryCount": 0  },  {    "user": "michael.abc@abc.com",    "Payload": "/content/mygps/us/en/pi/public/service/account-servicing/account-maintenance/restrictions-for-goal-based-planning-and-investing-umh-091021-mt_ditamap/Phonebooka17b4efc-716b-49cd-a909-480703ef2a74",    "topic": "com/abc/mygps/generatePDF",    "id": "2026/10/8/16/4/7564fa83-ecaf-4986-a1fb-da04e8eae38b_431",    "state": "QUEUED",    "queueName": "MyGPS Generate PDF Queue",    "createdTime": "2026-10-08T16:04:49.342Z",    "pendingTimeSeconds": 79082,    "pendingTimeMinutes": 1318,    "pendingTimeHours": "21 hr 58 Minutes",    "retryCount": 0  }

**********************

  

Best answer by social_portfolio81db

​@helloosuman ,

Hi,

Looking at your dump, I don’t think PDF generation itself is slow. It looks like your queue is blocked. Job _427 is ACTIVE and everything behind it is QUEUED with about 22 hours pending. Since the queue is ORDERED, it only runs one job at a time (maxparallel is ignored), so a single hung job at the front would hold up the rest.

Your pending time is calculated from createdTime, so it can’t tell you whether a job is waiting or actually running. I’d log job.getProcessingStarted() and the owning Sling ID as well. On Cloud Service, an ACTIVE job owned by an author pod that has since been replaced can sit there for a very long time.

As for why it’s happening after the AMS move, I’d look at two things first:

  • A call with no timeout somewhere in the consumer (Guides output, DITA-OT, or an external HTTP call). Egress is restricted on AEMaaCS, so something that worked fine on AMS can just hang.
  • Smaller author resources, which can lead to GC pressure or OOM during PDF generation.

To get things moving, I’d try these:

  1. Stop the active job with JobManager.stopJobById(). It only works if your consumer checks isStopped(). If it doesn’t, you’ll probably need Adobe support to get a thread dump.
  2. Switch to UNORDERED with queue.maxparallel around 2, if the PDFs don’t have to be generated in order.
  3. Add timeouts and start/finish duration logs in the consumer.

I can’t see your consumer code, so this is just my read of the dump and not a confirmed root cause. If you can share the aemerror lines around job _427 (around 16:04 UTC on Oct 8), the consumer logs, or the author CPU/memory from Cloud Manager for that window, I’m happy to take another look.

 

Regards,

Santhosh

2 replies

social_portfolio81dbAccepted solution
Level 2
October 9, 2026

​@helloosuman ,

Hi,

Looking at your dump, I don’t think PDF generation itself is slow. It looks like your queue is blocked. Job _427 is ACTIVE and everything behind it is QUEUED with about 22 hours pending. Since the queue is ORDERED, it only runs one job at a time (maxparallel is ignored), so a single hung job at the front would hold up the rest.

Your pending time is calculated from createdTime, so it can’t tell you whether a job is waiting or actually running. I’d log job.getProcessingStarted() and the owning Sling ID as well. On Cloud Service, an ACTIVE job owned by an author pod that has since been replaced can sit there for a very long time.

As for why it’s happening after the AMS move, I’d look at two things first:

  • A call with no timeout somewhere in the consumer (Guides output, DITA-OT, or an external HTTP call). Egress is restricted on AEMaaCS, so something that worked fine on AMS can just hang.
  • Smaller author resources, which can lead to GC pressure or OOM during PDF generation.

To get things moving, I’d try these:

  1. Stop the active job with JobManager.stopJobById(). It only works if your consumer checks isStopped(). If it doesn’t, you’ll probably need Adobe support to get a thread dump.
  2. Switch to UNORDERED with queue.maxparallel around 2, if the PDFs don’t have to be generated in order.
  3. Add timeouts and start/finish duration logs in the consumer.

I can’t see your consumer code, so this is just my read of the dump and not a confirmed root cause. If you can share the aemerror lines around job _427 (around 16:04 UTC on Oct 8), the consumer logs, or the author CPU/memory from Cloud Manager for that window, I’m happy to take another look.

 

Regards,

Santhosh

Level 2
October 10, 2026

@social_portfolio81db excellent analysis. As per your recommendation, let me implement the fix and keep all you posted about the solution.