Status lacked ownership
Officers described familiar labels as clear, then opened the case anyway to decide whether it needed them.
ITBA / Government services / 2024–Present
One assessment case, from arrival to order - and the four moments where an Income Tax Department officer had to work out whether it was still theirs.
Professional work Sanitised reconstruction
The same information, ranked by what an officer has to decide: what state, whose turn, how long left, one next step.
The central shift
From a status that names a state
to a status that explains who acts next.
02 / Problem & audience
ITBA is the internal application Income Tax Department officers use to manage cases, track statutory deadlines, and record decisions. This case follows a single assessment proceeding - arrival, notice, referral, order - because that one sequence is where the department's statutory deadlines, handoffs between offices, and record-keeping obligations all land on the same desk at the same time.
Officers knew the procedure. The interface still made them reconstruct case history, interpret generic status labels, and hunt for the action that mattered.
Make the state, risk, and next responsibility understandable without weakening the procedural record.
03 / Role & constraints
As a Senior UX Designer, my contribution covered research synthesis, information architecture, wireframes, prototypes, and testing materials. I worked within a team of three designers, with a PM, engineers, departmental subject-matter experts, and reviewing officers.
Fields and checks had to remain traceable to procedure. A simpler interface could not make a decision harder to review or defend.
No live screens, taxpayer data, or officer identities are shown. Interface examples on this page are illustrative reconstructions, not reproductions of the shipped system.
04 / Research & findings
Contextual inquiry, a procedural audit, and stakeholder workshops connected what officers did with what the process required. Prototype walkthroughs then tested the proposed structure.
Officers described familiar labels as clear, then opened the case anyway to decide whether it needed them.
Side notes and second windows held information that did not follow officers between related screens.
Personal pre-checks helped avoid errors that the interface surfaced too late.
Generalised observations from the study; these summaries are not participant quotations or measured production effects.
Four methods, each covering a blind spot in the others: officers could show where the workflow broke down but not why a field existed, and the audit could show why it existed but not whether anyone struggled with it.
| Method | Who took part | The question it carried |
|---|---|---|
| Contextual inquiry | IRS officers, ITOs, Inspectors, Office Superintendents, Tax Assistants and Data Entry Operators, at their own desks | Where does an officer hesitate, backtrack, or step outside the system? |
| Procedural audit | Field-by-field against departmental procedure, reviewed with SMEs | Which fields and steps are genuinely mandatory, and which are inherited? |
| Requirement workshops | Departmental SMEs and reviewing officers, with the PM | What would a proposed simplification cost downstream - in audit, handoff, or record-keeping? |
| Prototype walkthroughs | Participants from across the same six cadres, and departmental stakeholders | Does the redesigned flow hold up when someone is asked to work a case in it? |
Eight to twelve people over eight to twelve weeks - a range because sessions were fitted around live casework. A case passes through all six cadres above, so the study followed the chain rather than concentrating on the seat with the most authority, and sampled across proceeding types rather than seniority. Taxpayers are absent by design, and the downstream offices were represented through SMEs rather than observed.
Sitting with officers on live cases, rather than asking them to describe the process from memory, is the whole reason these findings exist: people do not reliably report friction they have already normalised. Three patterns recurred in nearly every session - described one way when asked, worked another way when watched. Paraphrased composites; no verbatim quotes or identities can be published.
Watched, the same officers opened the case anyway. Legible, but not decisive - it named a state without answering whose turn it was.
Watched, they scanned top to bottom, often twice, before an action taken many times before. Fluency was doing work the interface wasn't.
The same officers kept a personal pre-check - side notes, a second window, a colleague - to avoid it. The workaround had been normalised long enough to stop registering as friction.
An officer needs five things answered at any moment in a case. Every friction point from the sessions was logged against whichever one it obscured - and those five tags are the ones running down the left of the log below.
Sanitised here: case details, screen names, and identities removed, patterns kept.
Identity and proceeding were re-established by hand on each related screen, often by copying a reference into a side note first.
A second window stayed open purely to hold context the current screen dropped.
What had already happened on a case had to be reconstructed from separate records rather than read in one place.
Nothing distinguished a case with a statutory deadline approaching from one with months of runway.
Officers scanned the full field set before acting, even on a routine step repeated daily.
After an action was recorded, where the case went and who owned it next was inferred rather than stated.
Officers returned to cases they had already actioned, to confirm the action had taken.
Every field and dependency was mapped against departmental procedure one by one. It ran field-by-field because in a system whose records carry legal weight, "this seems unnecessary" is not grounds to remove anything - only a traceable link back to procedure is. Workshops with SMEs and reviewing officers then pressure-tested each simplification against its downstream cost in audit, handoff, and statutory record-keeping.
Grouping every friction point by which of the five questions it obscured left four problems recurring often enough to drive the direction - the four carried into Strategy & trade-offs, where each is paired with the decision it produced.
Recurrence alone didn't set the order. Each theme was weighed on three counts: frequency across sessions, procedural cost if left alone - a missed limitation date is not the same class of problem as an extra click - and whether a workaround already existed, since a normalised workaround means the system has quietly offloaded work onto the person. The four cleared all three; smaller findings cleared only the first and were logged rather than designed for.
The limits are worth stating as plainly as the findings, because they bound everything else on this page.
ITBA handles live taxpayer and departmental data, so specific screens and officer identities from this work can't be published - the findings above are described in generalised form, consistent with departmental confidentiality requirements.
05 / Evidence review
Fieldwork tells you what is happening at a desk this month. It does not tell you what has already been tried, answered, and quietly allowed to return. So the officer research ran alongside a documentary review of the public record from 2017 to 2026 - association correspondence, the system owner's published responses to field complaints, security circulars, and the 2026 transition material.
Two reasons it was worth the time. ITBA serves roughly 45,000 personnel across about 780 offices, and no realistic sample speaks for that spread - the written record reaches parts of it fieldwork never will. It is also the strand that can be published: everything here is public-domain, which is why it carries citations where the fieldwork can only carry paraphrase.
| What was reviewed | Why it is useful | Watch-out |
|---|---|---|
| Officers' association letters to the department head (Jan and Mar 2026) | Recent, detailed, first-hand problem lists naming specific failures | Advocacy documents, written partly to protect members from accountability - failures are over-represented |
| The system owner's published answers to 60 field complaints (Oct 2018) | Puts each complaint beside its official response, which is rare and very revealing | Eight years old and defensive by design; some issues may since have been fixed |
| Security circulars on login-token sharing (2019–2020) | Documents the gap between the rule and the behaviour, from the rule-maker's side | A policy view, not a user view |
| Federation bulletins, 2017 launch coverage, vendor material, new-Act transition guides | Hardware constraints, change history, scale, and the 2026 statutory shift | Background and context only; not evidence of any individual's experience |
Every documentary finding was tagged documented (stated directly in a source), inferred (strongly implied by it), or hypothesis (plausible from domain knowledge, unvalidated). Hypotheses were kept out of the recommendation set entirely - they became things to test, not things to act on. Only documented findings appear on this page.
This is the finding that changed the recommendation. Laid side by side, the 2018 record and the 2026 record describe the same categories of failure - and the 2018 responses show why.
| Complaint class | 2018 | 2026 | How it was answered in 2018 |
|---|---|---|---|
| System slowness | Officers losing hours waiting on the application | Extreme slowness and buffering | Storage expanded; declared resolved |
| Central processing delay | Days to months for a final order | Orders pending for months | Described as near real-time; treated as isolated cases |
| Deductions not allowed in computation | Deduction refused despite repeated entry | Still reported as sometimes not allowed | Attributed to user data entry; more training offered |
| Digital signature delays | Signing malfunction stalls the workflow | Signing delays of days | Installation instructions and field support |
| Management reports out of date | Slow and out of sync with the system | Not refreshed regularly | Monitoring and query optimisation |
| Tickets closed without a fix | Marked resolved without explanation | Closed without redressal | A separate ticket category added |
Every 2018 response added capacity, added training, or reframed the problem as user error. Not one of them made the system's internal state visible to the person accountable for the case. That is the mechanism by which the same complaints came back - and the reason this project argued against leading with a redesign.
The say-do gaps above came from watching officers work. These came from the documents - a harder kind of evidence, because each is a behaviour the organisation has written down about itself.
| Research finding | What changed as a result |
|---|---|
| Status didn’t answer the real question | Status states that mean something. Explicit states - incomplete, ready for review, awaiting another party, resolved - so progress is read at a glance, not inferred from fields. |
| The one action that mattered got lost in the density | One clear next action per screen. Primary actions, supporting information and secondary controls separated into a visible hierarchy. |
| Context didn’t travel with the officer | Case context that persists. Identity, proceeding and status stay available across related tasks instead of being rebuilt on every screen. |
| Validation caught mistakes too late, or not at all | Validation at the point of the mistake. Missing or inconsistent information surfaced next to the field, instead of failing a whole submission at the end. |
Read together with the observed gaps, the pattern is consistent enough to state as a principle: every one of these sits at a point where the system transfers cost to the user - time, uncertainty, or personal risk - and returns nothing. No status, no explanation, no control.
Collapsing the full symptom inventory produced six systemic causes. The column that matters most to me is the last one: a research recommendation that quietly implies design can fix infrastructure is a recommendation that will fail in front of an engineering lead.
Two of the six are the ones a disposal flow actually touches, so they are the two that shaped this work. The remaining four are real and were reported as found - they simply belong to the platform rather than to this case.
| Root cause | Mechanism | What design can change | What it cannot |
|---|---|---|---|
| Asynchronous black boxes | Signing, dispatch, central processing and schedulers all run in the background with no user-visible state | Status, receipts, expected durations, next actions, alerts on stuck items | Queue throughput and processing capacity |
| Computation trust deficit | Opaque rules produce mismatches, users bypass to manual upload, which creates more delay and less trust | Explainability, pre-submission validation, difference views | Backend rule defects |
These came out of the same documentary review and are listed for completeness. None of them is a case-disposal problem, and none of them shaped the four decisions on this page - a delegation model or a support SLA is not something an assessment screen can reach.
| Root cause | Mechanism | What design can change | What it cannot |
|---|---|---|---|
| A security model built for individuals | One person, one token, one session, in offices that work as teams | Self-serve delegation, role templates, a delegation audit trail | Security policy itself |
| Support that measures closure, not resolution | Tickets close without user confirmation; no shared view of known issues | In-context reporting, close-on-confirmation, duplicate clustering | Vendor service-level structure |
| Governance through stale reporting | Supervisors cannot trust the data, so they ask for manual reports, which costs casework time | Freshness timestamps, drill-down, report-once dashboards | Management culture and disposal targets |
| Infrastructure and change overload | Capacity, hardware and releases are not aligned to the statutory calendar, during simultaneous platform migrations | Status communication, change enablement, dual-regime assistance | Bandwidth, hardware provisioning, budgets |
Three directions were on the table: modernise the interface, make it faster, or add training. The documentary strand ruled out two on evidence rather than taste - training had been the 2018 answer and the complaints returned; speed was real but outside design's control. What was left, making the system's own state visible to whoever is accountable for it, was also the cheapest: the events were already being produced. That is why this project is about clarity, not a restyle.
06 / Strategy & trade-offs
Every decision below responds directly to one of the four research findings, rather than to a generic best practice.
| Research finding | What changed as a result |
|---|---|
| Status didn’t answer the real question | Status states that mean something |
| The one action that mattered got lost in the density | One clear next action per screen |
| Context didn’t travel with the officer | Case context that persists |
| Validation caught mistakes too late, or not at all | Validation moved to the point of the mistake |
The sharpest disagreement on this project was about density. The working cadres wanted less on screen - that is most of what the research found. Reviewing officers wanted fields kept visible, because a decision they may have to defend months later has to be reconstructable from the record. Each was correct about their own job, and a single default screen could only be shaped around one of them.
What resolved it was refusing to settle it on design authority. Nothing came off the screen because it looked cluttered; a removal had to trace to an actual procedural requirement, and anything that traced stayed regardless of how the screen read. That rule cost me one of my own decisions: a change to validation-message grouping was reverted once it turned out to make the reason for a flag harder to reconstruct - cleaner to read, worse to defend.
The lasting lesson was about who to ask. The people best placed to say what could be removed were not the ones using the screen most, and I had been weighting the cadres that touch the screen most, and the ones who inherit the consequences had the better answer about what could go.
07 / Workflow & wireframes
Low-fidelity flows mapped a single assessment proceeding end to end before interface design started. The sequence below is that case. Each of the four moments is one where the answer to is this still mine? changed - and where the old screen left the officer to work it out.
ArrivalThe case lands. Identity, proceeding and the limitation clock appear together, before any individual field competes for attention.
Notice issuedThe first action is recorded. The step just completed and the step now owed are both visible, so progress is read rather than inferred from fields.
Referred outThe case is waiting on another office. The state names the party actually holding it and how long it has been there - while the limitation date stays the officer's responsibility.
OrderEverything about to be submitted is shown together, with an explicit point to stop and reconsider.
08 / Testing & iteration
Wireframes and prototypes went back to participants from across the six cadres and to departmental stakeholders ahead of implementation - not for sign-off alone, but to test whether the model actually held.
Moderated, one officer at a time, on a clickable prototype. Tasks were framed as case work, not interface tasks - a case has come to you, take it as far as you can - so the officer chose the path rather than being walked down it. Nothing was explained in advance; where an officer asked what something meant, the question was the finding.
09 / Final experience
One assessment proceeding, shown at each of the four moments from the flow above. The Before column is the part worth watching: across all four it barely changes. The same field list, the same four equal buttons, the same word - Open - whether the case has just arrived or is twelve days from its limitation date. That is the finding, drawn rather than described.
ITBA holds live taxpayer data, so none of the real interface can be published. These are reconstructions - dummy references, invented case details, no departmental content. They exist to make the four decisions legible, not to reproduce what shipped.
Demonstrates case context that persists: identity, proceeding and the limitation clock in one place, before any field competes for attention.
The limitation date is not on this screen at all. Nothing marks the case as new to this officer, and four equal buttons give no indication that only one of them applies yet.
Identity, proceeding and the statutory clock arrive together, and the one step actually available is the only one emphasised.
Demonstrates status states that mean something: an explicit record of what completed and what is now owed.
The notice went out three days ago. The screen says what it said before it went out - “Open” - so whether the step registered has to be checked somewhere else.
The step just completed and the step now owed are both stated, so progress is read off the case rather than reconstructed from fields.
Demonstrates ownership in the status itself. This is the moment the old label was least useful, and the reason the project exists.
"Open" is accurate and useless. It doesn't say what is blocking the case, whose turn it is, or how much time is left - and four equal buttons don't say which one this case needs.
The same information, ranked by what an officer has to decide: what state, whose turn, how long left, one next step.
Demonstrates validation at the point of the mistake: the mismatch named where it occurred, rather than at the end of a submission.
The computation mismatch was present while the order was being drafted. Nothing said so until the submission came back rejected, with 12 days left on the limitation date.
The one unresolved item is named where it occurred and carries the corrective action. Submit stays available, but it is no longer the first thing the eye lands on.
Illustrative reconstructions only. No screen, field label, case reference, or record shown here is taken from the live system - the case numbers, dates, and masked references are invented for this page.
10 / Delivery & outcomes
The work produced a revised assessment flow, wireframes, prototypes, and testing materials. The evidence below comes from validation sessions and officer feedback; it does not establish a measured production improvement.
Delivery scope shown here: design and validation. The public case study does not specify rollout dates or implementation coverage. Operational time-to-resolution and escalation rates were not instrumented for a controlled before/after comparison.
| Capability | Before | After |
|---|---|---|
| Case status | A generic open / closed label | Explicit states - incomplete, ready for review, awaiting another party, resolved |
| Primary action | Buried among a dozen equally-weighted fields | Visually distinct from supporting detail and secondary controls |
| Case context | Rebuilt manually on every related screen | Persists - identity, proceeding, and status stay visible |
| Validation | Surfaced only at final submission, if at all | Flagged next to the exact field, before submission |
An officer could tell whether a case needed attention without opening it. In validation sessions, officers were noticeably faster to identify blocked cases from the state alone - the clearest behavioural signal from this project, though not something instrumented at scale afterward.
Separating the primary action from supporting fields was meant to remove the need to hunt across a dense screen for what to do next - officers confirmed that in walkthroughs, though production usage wasn't tracked to confirm it held at scale.
Catching a mismatch next to its source, before submission, was the change officers pointed to most often as removing rework - the strongest qualitative signal from this project, even without a formally measured rework rate.
No production metric was captured. That is a gap I can at least be specific about: each measure below tracks a behaviour this study watched fail, and would move only if one of the four design decisions worked.
That rules out the measures it would be most tempting to claim - dispatch time, processing delay, report freshness. The root-cause table above assigns those to queue throughput and management culture, which interface work does not control. A measure that cannot fail because of my work is not evidence for it.
| Measure | The decision it tests | What it would actually show |
|---|---|---|
| Share of cases correctly triaged from the worklist without being opened | Status states that mean something | Whether a state label now answers "is this one mine?". Officers were observed opening cases purely to find that out. |
| Return visits to a case already actioned, per officer per week | Status states that mean something | The log recorded officers going back to confirm an action had registered. If the status says so, this should fall - and it needs no new event, only a count of what is already written down. |
| Whether an officer's first action on a case is the one the case needed | One clear next action per screen | Whether the primary action is found rather than arrived at after a full scan. Time-to-first-action is the cheap version; whether that first action is correct is the one worth having. |
| Screens opened per case, and concurrent sessions per officer | Case context that persists | A second window stayed open purely to hold context the current screen dropped. If context travels with the officer, that window has no job left. |
| Validation flags corrected in place, as a share of all flags raised | Validation at the point of the mistake | Whether a flag is now actionable where it appears, rather than sending an officer to another screen, a side note, or a colleague. |
One system-level measure belongs alongside these without belonging to this work: the share of orders pushed through manual upload. It needs no new instrumentation, and it converts a question about feelings into one about what officers do when nobody is asking. It tracks the computation-trust deficit rather than this interface, so it is not a claim I would make for this design - but it is the number I would fight hardest to have collected.
11 / Reflection & next steps
The rules, exceptions, and accountability requirements inside ITBA are real; removing them was never the brief. The work was to make that complexity navigable - to surface the right information at the right moment, and let officers act with more confidence on decisions that carry real legal weight.
What travels beyond assessment is the frame, not the screens. The five orientation questions are what any officer needs answered in any ITBA module, and the status vocabulary and in-context validation were built to be reused that way. The fields, statutory checks and sequence are assessment's own, and would have to be re-derived for appeals or exemption rather than copied across.
Observe downstream roles directly, rather than relying only on their representation through SMEs and reviewing officers.
Instrument time-to-resolution and escalation rate to understand whether the observed benefits hold in everyday use.