Blog

11 Things Your SOX Program Is Probably Doing the Hard Way

Subscribe now to join the Risk Register community:

Almost nobody made a bad decision to get here.

That's the thing about an inefficient SOX program. It's rarely the product of a mistake. It's a sequence of entirely reasonable choices that calcified into an annual routine, and then nobody had a spare month to question the routine while they were busy executing it.

The first ten items below recur most often, cost the least to fix, and produce the biggest drop in effort per unit of assurance. None of them reduce coverage. None require you to buy anything. Most can be put in place inside a single testing cycle, and several can happen before this year's fieldwork starts.

The eleventh is different in kind, and it's the one you're most likely to be asked about before the year ends.

Recognize four or five of these and your program is normal. Recognize nine and next year could be a great deal easier than this one.

1. You Walk Through the Control, Then Test It All Over Again

Design evaluation usually happens early in the year through walkthroughs. Operating effectiveness testing happens later, as a separate exercise, and often opens by re-establishing exactly the understanding the walkthrough already produced. Same control owner, same process, same conversation, four months apart.

Structure the walkthrough to include an examination of one real transaction, documented to testing standard rather than to note-taking standard. You confirm design and produce your first piece of operating evidence in one session.

One meeting instead of two, one request instead of two, and your design conclusion rests on an actual transaction rather than a description of one. Whether that transaction also counts toward your operating effectiveness sample depends on documentation quality and period coverage, so agree the treatment with your external auditor rather than assuming it.

2. Everything Gets Tested at Year End

This is the most common structural inefficiency in SOX, and the most consequential.

Testing compresses into the fourth quarter. Control owners get asked for a year of evidence at the precise moment they're closing the year. And every exception surfaces when there's no time left to remediate it, retest it, or do anything about it except argue over severity.

Test at interim instead, typically through the third quarter, then run roll-forward procedures for the remaining period. Roll-forward is a fraction of the effort of the original test: confirm nothing changed in design, ownership, or system configuration, then test a small number of samples from the stub period.

The workload spreads across the year rather than colliding with close. Exceptions surface in month eight, when remediating and retesting inside the same fiscal year is still genuinely possible. And your control owners get asked for evidence during a quarter when they can actually produce it, which is worth more to the relationship than most program changes you could make.

3. You Retest Automated Controls From Scratch Every Year

An automated application control that hasn't changed, running in a system whose change management and access controls are effective, produces the same result this year that it produced last year. Plenty of programs test it fully anyway, every year, out of habit.

AS 2201 permits a benchmarking strategy for automated application controls in subsequent years' audits, and paragraphs .B28 through .B33 set out how it works. The reasoning at .B28 is that entirely automated controls generally aren't subject to breakdowns from human failure, so once you've established a baseline, you confirm no change occurred rather than retesting the logic.

Automated controls are already the cheapest controls in your population. Benchmarking makes them cheaper still.

One thing the shorthand version of this advice usually leaves out, and it matters. Benchmarking isn't permanent. Paragraph .B33 provides that the baseline should be reestablished after a period of time, with the auditor weighing the effectiveness of your IT control environment, the nature of any changes to the programs containing the controls, the nature and timing of other related tests, the consequences of errors associated with the benchmarked control, and whether the control is sensitive to business factors that may have shifted. So you need effective ITGCs, a reliable change log, and a view on when the baseline gets refreshed. Agree the approach with your external auditor before you rely on it.

4. Every Sample Triggers Its Own Population Request

You know the pattern. Request evidence for sample one, receive it, request sample two, receive it, repeat until somebody loses patience. Twenty samples become twenty round trips with the same control owner. And in most versions of this, nobody ever validates that the population the samples came from was complete.

Request the whole population once. Test it for completeness and accuracy. Then select and request every sample in a single batch, from a population you've already validated.

One request cycle instead of twenty, and you settle the completeness question at the start rather than discovering during review that the population was a filtered extract and the entire selection has to be redone.

5. You Test the Same Report Six Times

Information produced by the entity, meaning the reports, queries, and system extracts your control operators rely on, has to be tested for completeness and accuracy. A control based on an unreliable report isn't a reliable control.

In most programs the same aging report or GL extract supports several different controls, and each tester independently tests it again. Three people, three sets of procedures, one report.

Build an inventory of the reports your controls depend on. Test each one once, at the right level of rigor, and reference that testing from every control that relies on it.

Report testing is unglamorous and slow, so removing duplicated effort is a direct saving. It also fixes a quieter problem, which is three testers reaching three slightly different conclusions about the same report and nobody noticing until review. Incomplete or untested IPE is among the most common reasons external auditors send workpapers back, so consolidating it raises quality at the same time it cuts hours.

6. Your Management Review Controls Have No Defined Precision

Management review controls are the hardest controls in a SOX population to test and the most frequently challenged, and the reason is almost always identical. The control is documented as "the controller reviews the account reconciliation," which describes an activity rather than a control.

Define three things for every review control before testing starts. What threshold triggers investigation. What the reviewer is expected to do when something exceeds it. And what evidence demonstrates the investigation happened and the item got resolved.

Without precision criteria a tester can't conclude whether the control operated, because there's no standard to evaluate it against. So testers document that a review occurred, auditors challenge that as insufficient, and the control gets retested with better criteria inside the same year. Defining precision up front skips the entire loop.

It also does something uncomfortable and useful. It regularly reveals that a review control everyone assumed was strong has no threshold at all, and never did.

7. You Send Control Owners Requests Instead of Specifications

"Please send evidence that the reconciliation was reviewed" produces whatever the control owner believes that means. Which produces a follow-up. Which produces another one.

Write an evidence specification for each control, once: exactly which document, which fields have to be visible, what constitutes evidence of review, what format is acceptable. Send the specification with the request. Reuse it every year after that.

Control owner time is the most expensive time in a SOX program, because it isn't on your budget, it's on theirs. It's also what determines whether the business experiences SOX as a partnership or a tax. Specifications collapse three exchanges into one, and they make your program portable, because a new tester can execute the control without inheriting institutional knowledge nobody wrote down.

8. Your Samples and the Auditor's Samples Are Different

Your team selects samples. Your external auditor selects different samples, from the same populations, for the same controls. Your control owners produce evidence twice and quietly conclude that nobody is talking to anybody.

Agree scope, timing, sample sizes, and selection methodology with your external auditor before fieldwork begins, so the auditor can reuse your selections where they intend to place reliance.

This is the most visible improvement you can make to the business, because the duplicate request is the specific thing control owners complain about. It also reduces the auditor's own hours where reliance is placed, which is a fee conversation worth having explicitly rather than hoping shows up.

Set expectations honestly, though. Reliance isn't automatic. Under AS 2201 it depends on the competence and objectivity of your team and the quality of your workpapers, and it's capped by risk, because as the risk associated with a control rises, so does the auditor's need to test it themselves. Reliance builds over two to three years. It doesn't arrive in one.

9. Deficiency Evaluation Happens in December

Exceptions get logged during testing and evaluated at the end, in aggregate, under time pressure. Which means severity conclusions get argued with your auditor at the exact point in the year when there's no remaining opportunity to remediate and retest.

Evaluate every exception when you find it. Root cause, isolated or systemic, magnitude, likelihood, compensating controls. Keep a running aggregation analysis rather than assembling one in a weekend.

A deficiency found in month seven can often be remediated and the control retested inside the same fiscal year, which changes your year-end conclusion. The same deficiency found in month eleven cannot. Continuous evaluation also means you make the severity argument once, with evidence, instead of opening a batch negotiation with an auditor who has already formed a view.

10. You Re-litigate the Same Scoping Questions Every Year

Why is this entity out of scope. Why is this account below the threshold. Why is this control non-key. Why was this location excluded. Every year the same questions come round, sometimes from a new audit team, and every year somebody reconstructs the answer from memory and a spreadsheet they think is the right version.

Keep a scoping decision log. One line per decision: what was decided, the quantitative or qualitative basis, who approved it, when, and whether the external auditor concurred. Update it during scoping rather than reassembling it under questioning.

Institutional memory stops depending on individuals, which matters more than it sounds like it does when the person who made the call has moved on. A new auditor gets a documented rationale instead of somebody's recollection, which is both faster and considerably more credible. And when a decision genuinely needs to change because materiality moved or the business did, you can see what the original basis was and whether it still holds.

11. You Don't Have an AI Position, So You Have Eleven Unofficial Ones

The first ten are routines to fix. This one is a gap to close, and it's the item most likely to arrive as a question from your audit committee, your external auditor, or your CFO before the year is out.

Almost no SOX program has a written position on AI. Nearly every SOX program already has AI in it.

Somebody on the team drafted a process narrative with a chatbot. Somebody pasted a population extract into a tool to find duplicates. Somebody used an assistant to summarize a control owner's email thread into a workpaper note. None of them did anything unreasonable, and none of them asked, because there was nothing to ask against.

That leaves you exposed in two directions at once, which is what makes it worth handling now rather than next cycle.

The first exposure is adoption you can't explain. Your workpapers support a certification your CEO and CFO sign personally. If AI contributed to a workpaper, you need to be able to say how, what data went into it, and that a competent person reviewed the output and owns the conclusion. "I'm not sure which tool that was drafted in" is not a sentence you want to say during fieldwork. Separately from SOX entirely, company financial data flowing into consumer tools is a data handling question somebody should already be asking.

The second exposure is the work you're still doing by hand. While the ad hoc adoption creates risk, the program continues to run on manual evidence chasing, spreadsheet trackers, and sampling in places where full populations are sitting right there. The absence of a strategy isn't caution. It delivers the worst of both outcomes: unmanaged use at the margins, and no benefit where the benefit is real.

Write a one-page position answering four questions, and circulate it before your next testing cycle.

What may AI touch? Draft narratives, summarize documents, analyze populations, suggest test attributes. Be explicit about what it may not touch: concluding on operating effectiveness, evaluating deficiency severity, or any judgment supporting the assertion.

What data may leave your environment? Name the approved tools. State whether company data may be entered, and whether that data can be used for model training. If the answer is no, get it in the vendor terms rather than on a marketing page.

How is AI-assisted work documented? Provenance and review. Which steps involved AI, and evidence that a qualified person reviewed the output. This is the piece almost nobody has, and it's the first thing an auditor will ask about.

Who approves a new tool? One named owner. Without one, approval defaults to whoever downloaded it.

As for where AI earns its place in a SOX program today, it's in evidence collection, annotation, and status tracking, which is where the manual burden actually sits. Full population analysis in place of sampling where the data supports it. Reproducible recomputation an external auditor can re-run themselves. First drafts of process narratives and risk and control matrices for human review. Triage of exceptions, so judgment work reaches the people who should be doing it.

Where it doesn't belong is conclusions, severity evaluations, and anything the certification rests on. The technology handles the mechanical work and the practitioner owns the judgment. Full stop.

A written position turns an unmanaged risk into a governed capability. It gives your team permission to use the tools that save real hours, boundaries that keep company data where it belongs, and a documentation standard that survives questioning. It also means that when the audit committee asks what your position on AI is, you have one. Broader AI governance is its own discipline with its own frameworks. Explore our AI governance and emerging risk services.

Where to Start If You Only Pick Three

Take number 2, interim testing with roll-forward. It has the largest effect on workload and it changes the character of the program from a year-end event into something managed.

Take number 7, evidence specifications. Cheapest to implement, most visible to the business, and it compounds every year afterward.

Take number 6, review control precision. Highest quality return, because management review controls are where external auditors challenge most often and where rework costs the most.

Two of these are timing-sensitive, so they're worth putting on a calendar rather than a list. Interim testing has to be planned before the testing calendar is set. Sample coordination with your external auditor has to happen before fieldwork. Both windows close early in the year and neither reopens.

Number 11 sits outside that ranking because it isn't an efficiency, it's an exposure. If you have no written AI position and your team has access to AI tools, close that before optimizing anything else. It takes an afternoon, and it's the only item here where waiting costs you risk rather than hours.

One structural note to finish on. Every item above makes an existing population cheaper to test. Not one of them asks whether the population is the right size in the first place. Before optimizing how you test a hundred and fifty controls, it's worth asking whether the answer is a hundred and fifty. More on control rationalization.

The routine will keep calcifying whether or not anybody looks at it, and next year's version will feel just as normal as this one does. The programs that stay efficient aren't the ones that found a better tool. They're the ones where somebody is asked, once a year, what we're still doing the hard way.

Ten of these need somebody with the time to look at the routine instead of executing it. The eleventh needs a decision. If either one is your constraint, explore our SOX and internal controls practice, or reach out and we'll work out which three are worth your next quarter.

Cherry Hill Advisory is a global practitioner-built internal audit and risk advisory firm, led by former CAEs and Big Four alumni, delivering co-sourced internal audit, EQA conformance, SOC 2 readiness, ERM, fraud risk management, SOX compliance, cybersecurity, and AI governance. IIA Authorized Licensee. NASBA-accredited CPE provider.

Frequently Asked Questions

What is roll-forward testing in SOX?

Roll-forward testing extends an interim testing conclusion through the rest of the period. It confirms that the control's design, owner, and system configuration didn't change, and tests a limited number of additional samples from the period after interim testing. It takes substantially less effort than testing the full period at year end.

What is benchmarking of automated controls?

Benchmarking lets an automated application control tested in a prior period be relied on in subsequent periods without full retesting, provided the control hasn't been modified and the general controls over program change and access are operating effectively. PCAOB AS 2201 recognizes the approach at paragraph .60 and describes it at .B28 through .B33, including the requirement to reestablish the baseline after a period of time. Agree the specific application with your external auditor.

Why do external auditors challenge management review controls so often?

Because most review controls are documented as activities rather than as controls. Without a defined threshold that triggers investigation, a defined expectation for what the reviewer does, and evidence that outliers were investigated and resolved, there's no standard against which operating effectiveness can be concluded. Defining precision criteria before testing resolves most of these challenges.

What is IPE in SOX testing?

IPE stands for information produced by the entity: the reports, queries, and system extracts a control operator relies on to perform a control. IPE has to be tested for completeness and accuracy, because a control built on an unreliable report isn't a reliable control. Testing each report once and referencing it across every control that relies on it removes the duplicated effort.

How can we reduce the burden SOX places on control owners?

Three changes produce most of the improvement. Write an evidence specification for each control instead of sending an open-ended request. Request complete populations once rather than sample by sample. And coordinate sample selections with your external auditor so owners aren't asked for the same evidence twice.

Can these improvements be made in the middle of a testing cycle?

Some can. Evidence specifications, population request discipline, IPE consolidation, review control precision, continuous deficiency evaluation, and a written AI position can all go in mid-cycle. Interim testing with roll-forward, benchmarking, and sample coordination with the external auditor get planned before the cycle starts, so they apply to the following year.

Can AI be used in SOX testing?

Yes, for mechanical work: evidence collection and annotation, population analysis, reproducible recomputation, exception triage, and first drafts of narratives and control matrices for human review. It shouldn't be used to conclude on operating effectiveness or to evaluate deficiency severity, because those judgments support management's assertion and need a qualified person to own them.

How should AI-assisted work be documented for the external auditor?

Provenance and review. Record which steps involved AI assistance, what inputs were used, and evidence that a qualified person reviewed the output and owns the conclusion. Auditors are increasingly asking how a workpaper was produced, and a program that can't answer will be asked to redo the work.

What should a SOX program's AI position cover?

Four things at minimum: what AI is permitted to touch and what it isn't, what data may be entered into which tools and whether that data can be used for model training, how AI-assisted work is documented and reviewed, and who approves adoption of a new tool. One page is enough. The absence of a position doesn't prevent adoption, it just makes adoption unmanaged.

Subscribe now to join the Risk Register community:

Nk it'