Data changelog

The searchable, centralised archive of Australian FOI releases

↑ Data & downloads

This file records changes to data that had already been published — entries altered, removed, relabelled, or reinterpreted after they first appeared in a public file.

It does not record routine collection. New entries arriving because an agency added them to its disclosure log, and new agencies being added to the project’s coverage, are the normal business of the dataset and are not listed here except where a coverage change also affected existing entries.

This changelog covers the published artefacts: foiforest.csv and the per-agency raw CSVs at data.foiforest.org.

retired_ids.csv is a companion to this changelog, not a dataset: it lists identifiers that stopped resolving, and it carries no coverage guarantee of its own.

foi_data.json and agencies.json are also fetchable from that domain, but they are generated to drive the website. They are not versioned or documented as products, their shape may change without notice, and this changelog does not cover them. If you want the data, take the CSV.

Entries cite a seven-character commit hash such as a1b2c3d. The project’s source code is not public, so these are not links and cannot be looked up — they are included so that each entry is anchored to one specific change, and so that they resolve if the source is ever opened. Everything an entry actually asserts is stated in the entry itself.

About versions and reproducibility

No historical versions of the published files are available. Every hour from 8am to 6pm Melbourne time, Monday to Friday, the project checks agencies’ disclosure logs for new entries. When it finds any, it rebuilds the published files and overwrites the previous build at the same URLs. The files are also rebuilt once a week after a full check of every log, and occasionally when a correction is made. There is no version number, no dated snapshot, and no archive of previous builds.

On most weekdays the files change several times. A date does not identify a copy: two downloads on the same day can differ, and the differences are not always additions. Changes that alter entries already published are recorded in the entries below.

If you use this data for anything that needs to be reproducible, such as research, a published figure, or a claim someone might check, download your own copy and record two HTTP response headers from the download:

  • Last-Modified: when the file was uploaded. Cite this time, not only the date.
  • ETag: a checksum of the file’s content. Two downloads with the same ETag are the same file. A build can upload an unchanged file again, so Last-Modified can change while the ETag stays the same.

curl -sI https://data.foiforest.org/foiforest.csv prints both.

Prior builds are retained privately in the project’s own storage. Every build that changed the dataset is kept as a separate version: 553 between 5 April and 28 September 2026. Nothing has been deleted. These versions are not published, but they are why the reconstructed entries below could be measured rather than guessed at.

On request_id

request_id is stable: the same entry keeps the same identifier from build to build. Since 28 September 2026, it is also unique. Every row in foiforest.csv has its own request_id, and you can use it as a key.

Identifiers are derived from content: an entry’s agency and its reference number, or, where there is no reference number, a hash of the agency, date and description. Where an agency gives more than one entry the same reference number, one entry keeps that identifier. Each other entry gets the identifier followed by its release date, for example Department of Veterans' Affairs:77505:2026-09-05, and a six-character code where two entries share a date.

A change to the content an identifier is derived from produces a new identifier for the same entry, which is what happened in several of the changes below. Identifiers that stopped resolving are listed in retired_ids.csv, mapped to their current identifier where that could be established.

Before 28 September 2026, request_id was not unique and was not a primary key. Entries to which an agency gave the same reference number shared one identifier. As at 3 August 2026, 134 identifiers were each shared by more than one entry, covering 286 rows. In every case the entries were genuinely different FOI requests: Home Affairs FA 18/12/00916 covered seven separate requests, and the largest concentrations were Home Affairs (48 rows), Health (46) and ASIC (24). If you hold a copy from before that date, request_id alone does not identify a row in it. In the 3 August 2026 build, no two of those rows shared both a description and a date, so request_id, description and date together told them apart.


Reading retired_ids.csv

retired_ids.csv lists every identifier that has stopped resolving. One row per retired identifier:

Column Meaning
retired_id the identifier that no longer resolves
current_id the entry’s identifier today, where it could be established
resolution see below
change_id a slug naming the change that retired it, matching an entry in this changelog
date_retired the date the change was made
reason short description of the change

resolution takes three values:

  • mapped: the entry is still in the dataset, and current_id gives its identifier today.
  • unmatched: no single current entry could be matched to the identifier, because the entry’s description or date changed at the same time. current_id is empty.
  • removed: the entry’s content is no longer in the dataset, and it has no successor. current_id is empty. This means the content is gone, not merely that a row was dropped. Rows are frequently dropped because the same entry is held from another source, and those are mapped or unmatched, never removed. Three identifiers carry this value: the three empty Australian Electoral Commission entries removed on 28 August 2026.

What a retirement says about the entry depends on the kind of change that retired it:

  • Duplicates collapsed (2026-07-21 both entries, 2026-08-09, 2026-09-24, 2026-09-27, and the two 2026-09-28 Health entries): the entry is still in the dataset by construction, because a collapse keeps the surviving copy. Each identifier is mapped, or unmatched where only the automatic match failed.
  • Records reassigned (2026-04-24 CASA, 2026-04-25 successor departments, 2026-08-14 file sizes): the entry is still in the dataset by construction, because a reassignment relabels an entry rather than dropping it. Each identifier is mapped, or unmatched where only the automatic match failed.
  • A source removed (2026-07-24 Treasury): there is no such guarantee. Those entries were expected to be held from another source, but that could only be confirmed for 8 of 217. Each identifier is mapped or unmatched.
  • Entries removed (2026-08-28 Australian Electoral Commission): the content is gone. Each identifier is removed.

Entries

2026-09-28 — Two DAFF entries return to DAFF, and three descriptions follow the agency’s latest wording (ef6dde3)

DAFF rows are published under DCCEEW where DCCEEW holds the same request, because DAFF’s log still carries records from the former Department of Agriculture, Water and the Environment. The match used the reference number alone. Two DAFF entries shared a number with unrelated Department of the Environment entries from 2012 and 2014, and were published under DCCEEW. A shared number now counts only when the DCCEEW entry is dated within two years, or names the number in its description.

Separately, three agencies removed a final full stop from an entry’s description, and the dataset kept the older wording. It now publishes the latest wording, as for the six descriptions changed earlier today.

No entries added or removed. No identifiers retired or added.

What a consumer would observe. In foiforest.csv, five changed rows, and no other field changes:

  • DAFF 190112 (14 February 2019) and 200114 (23 March 2020): agency changes from “Department of Climate Change, Energy, the Environment and Water” to “Department of Agriculture, Fisheries and Forestry”. Their request_ids already carried the DAFF name and do not change.
  • Commonwealth Ombudsman FOI-2026-80017, Defence 533/24/25 and Industry 26/044/300368M: the description loses its final full stop.

No entry is added to the Atom feeds. The per-agency raw CSVs are unchanged.

Which reuse pattern breaks. Anyone counting or filtering entries by agency: DAFF gains two entries and DCCEEW loses two.

2026-09-28 — Six descriptions now follow the agency’s latest wording (4c70ce7)

Where an agency rewrites an entry’s description, the dataset publishes the latest version it has seen. A safeguard kept the older version instead when the newer one was less than half its length, because an IP Australia scraping fault had once cut descriptions short. That fault is fixed, and the safeguard is removed. Five descriptions now follow the agency’s own shorter rewrite, and one IP Australia description loses text the agency has removed from its page.

No entries added or removed. No identifiers retired or added.

What a consumer would observe. In foiforest.csv, six changed descriptions, and no other field changes:

  • Defence 0086/24/25: “Take a Closer Look Campaign For access to the documents” → “Take a Closer Look Campaign”
  • Defence 872/23/24: “Logistics Services For access to the document” → “Logistics Services”
  • NDIS Quality and Safeguards Commission 118 – 25/26 – (3): the longer summary is replaced by the Commission’s current one, “For the period 1 June 2025 to 31 January 2026: Total headcount per each Division from monthly report …”
  • National Anti-Corruption Commission FOI 26/73: the request loses its opening (“Dear National Anti-Corruption Commission, …”) and begins “I seek access to all documents …”
  • National Disability Insurance Agency FOI 24/25-0011: the longer summary is replaced by “Statistics around the conditions Borderline personality disorder (BPD) and Post traumatic stress disorder (PTSD).”
  • IP Australia 23/7056: the request text after the title is removed, leaving “Decisions relating to Essentially Derived Varieties – Plant Breeder’s Rights”

No entry is added to the Atom feeds. The per-agency raw CSVs are unchanged.

Which reuse pattern breaks. Anyone matching these six entries on description text.

2026-09-28 — IP Australia: two descriptions regain the request text (3929ce5)

Two IP Australia entries were published with only their title as the description. For 23/7450, IP Australia’s page changed its layout in May 2026, and the request text was no longer read. For 24/10111, the request text sits in an accordion that also holds 24/9991, and was never assigned to it. Both now carry the request text.

No entries added or removed. No identifiers retired or added. IP Australia only: 23/7450 and 24/10111.

What a consumer would observe. In foiforest.csv, two changed descriptions. 23/7450 changes from “The appointment of the Commissioner of Patents, …” to “Copies of documents relating to the appointment of the Commissioner of Patents, …”. 24/10111 changes from “Indigenous Liaison Officers” to “Documents containing information about Indigenous Liaison Officers or equivalent. ; The number of IP Australia staff who identify as Aboriginal or Torres Strait Islander.”. No other field changes, and no entry is added to the Atom feeds. In raw/ipa.csv, 3 rows have a new description and row_hash: the two above, and 1108, whose published description does not change.

Which reuse pattern breaks. Anyone matching the raw IP Australia file on row_hash: those 3 hashes change.

2026-09-28 — Every published row has its own identifier (18c22db)

A request_id is the agency and the reference number, so rows sharing a reference number shared one identifier: 123 identifiers were each held by more than one row, 264 rows in all. Most are distinct requests an agency lists under one reference, such as Home Affairs FA 18/12/00916, and separate releases under one request. Each row now has its own identifier. The row that held the identifier first keeps it. Each other row gets the identifier followed by its release date, for example Department of Veterans' Affairs:77505:2026-09-05, and a six-character code where two rows share a date.

No entries added or removed. No identifiers retired; 141 identifiers added. Every existing identifier still resolves, to the row first collected under it. By agency: Home Affairs 27, ACIC 12, ASIC 12, Health 12, ABC 10, Infrastructure 10, Defence 5, NBN Co 5, 23 other agencies 48.

What a consumer would observe. In foiforest.csv, 141 rows show a new request_id, and every request_id value is unique. No other field changes. Seven of those rows were already in the Atom feeds under the shared identifier (DVA 3, ARC 2, Finance 1, APSC 1). They leave the feeds rather than reappear as new entries. The per-agency raw CSVs are unchanged.

Which reuse pattern breaks. Anyone who looked up a shared identifier and expected every row: it now returns one row, and the others have new identifiers. Anyone who treated request_id as not unique can now use it as a key.

2026-09-28 — 27 Health entries published twice are merged where a year followed the reference’s letter code (18c22db)

Health lists some entries under two renderings of the reference number: its archived pages add a year after the letter code (25-0013 LD-2025), where the live log has none (FOI 25-0013 LD), or the other way round. These were published as two rows. They now merge, and the live entry keeps its identifier. Where both renderings carry a year and the years differ, the rows stay apart: 26-1965 MO-2025 and MO-2026, and 25-0019 LD-2025 and LD-2024, are different requests.

27 entries removed, none added. 27 identifiers retired, none added. Health only. Each retired identifier maps to one surviving entry with the same release date; the mapping is in retired_ids.csv.

What a consumer would observe. In foiforest.csv, 27 fewer Health rows. One surviving row changes: 25-0012 LD gains the document of its renamed copy 25-0012LD-2024, and its date_first_scraped moves from 2026-07-20 to 2026-04-05, when the renamed copy was first collected. No entry is added to the Atom feeds. The per-agency raw CSVs are unchanged.

Which reuse pattern breaks. Anyone joining on request_id: the 27 identifiers stop resolving, each with one successor, the same entry under the reference the live log uses.

2026-09-28 — 6 Health entries published twice are merged by a reviewed list (03ab1c1)

Health published six entries twice under two references the automatic merge does not match. One entry was renumbered by Health, from FOI 26-1965 MO-2025 to FOI 26-2028 MO-2025. Five are archived copies whose reference carries an earlier suffix than the live entry with the same number and release date. Each pair was checked against Health’s live log and is now merged by a reviewed list. The live entry keeps its identifier, its fields and its first-seen date. The other row’s documents are added to it.

6 entries removed, none added. 6 identifiers retired, none added. Health only. Retired, with the identifier that replaces each:

  • FOI 26-1965 MO-2025 (released 16 October 2025) → FOI 26-2028 MO-2025
  • 25-0197 LD-2025 → FOI 25-0197 LD IR
  • 25-0202 → FOI 25-0202 LD MO
  • 25-0255 LD-2025 → FOI 25-0255 LD MO
  • FOI 25-0275 → FOI 25-0275 LD
  • MO FOI 25-0114 LD-2025 → FOI 25-0114 LD MO

The Health entry FOI 26-1965 MO-2025 released 13 October 2025 and 26-1965 MO-2026 are different requests and stay as they are. The mapping is in retired_ids.csv.

What a consumer would observe. In foiforest.csv, 6 fewer Health rows. No surviving row changes: each merged-away row’s documents were already on the live row or it had none. No first-seen date moves and no entry is added to the Atom feeds. The per-agency raw CSVs are unchanged.

Which reuse pattern breaks. Anyone joining on request_id: the 6 identifiers above stop resolving. Each has one successor, the same entry under the reference Health lists today.

2026-09-28 — 57 entries regain their date of access (7e090b7)

When two collected copies of one entry are merged, the published row takes its dates from one copy. Where that copy had no date of access and the other did, the date of access was dropped. This happened to six ASIC entries after ASIC’s newest table stopped showing a “Date of access” column on 27 September, and to older Education and eSafety entries held from two sources. A merged row now keeps the date of access that any of its copies carries.

No entries added or removed. No identifiers retired or added. Rows changed by agency: eSafety 26 (LOG-3 to LOG-28), Education 25, ASIC 6 (FOI 045-2026, FOI 046-2026, FOI 047-2026, FOI 120-2026, FOI 121-2026, FOI 124-2026).

What a consumer would observe. In foiforest.csv, 57 rows where date_of_access changes from empty to a date, and date_source from date_published to date_of_access. In every one the date of access is the same day as the date published, so date, date_precision and date_basis do not change. No first-seen date moves and no entry is added to the Atom feeds. The per-agency raw CSVs are unchanged.

Which reuse pattern breaks. Anyone who used an empty date_of_access, or date_source of date_published, to find entries whose agency did not publish an access date: these 57 no longer match.

2026-09-27 — 36 entries published twice are merged, and 17 descriptions regain their punctuation (76cc9d5)

Some agencies list one entry twice: an archived copy beside the live one, the old Health website beside the new, or a Fair Work Commission entry dated by its financial year beside the same entry with its release date. These were published as separate rows. Two rows sharing a reference number are now merged when they have the same release date (or one is undated, or carries an FWC financial-year date) and substantially the same description, unless their descriptions name different identifiers such as grant or contract numbers. Four Health entries whose old and new website descriptions share too few words are merged by a reviewed list. Where an agency’s live log lists an entry twice with different descriptions, the two stay separate rows, as they are on the agency’s log.

Some archived pages stored punctuation in a broken encoding. Part of each broken character was removed on publication, so Treasury’s “Treasurer’s” was published as “Treasurers” and the Bureau of Meteorology’s “Bureau’s” as “Bureauâs”. The punctuation is restored.

36 entries removed, none added. 19 identifiers retired, none added. By agency: Health −19, Fair Work Commission −7, Services Australia −2, ART, ASQA, DAFF, DEWR, DSS, NHMRC, OAIC and TGA −1 each. Every retired identifier maps to one surviving entry; the mapping is in retired_ids.csv.

What a consumer would observe. In foiforest.csv, 36 fewer rows and 22 changed descriptions: 17 regain an apostrophe, dash or quotation mark (Treasury 7, Fair Work Commission 4, Bureau of Meteorology 2, Federal Court 2, Defence 1, AIHW 1), and 5 show the version most recently published by the agency (for example Defence “DIGGERWORKS” is now “Diggerworks”). Three Fair Work Commission entries change date: 19/20-26, 19/20-40 and 19/20-45 were dated 1 July 2019, the start of the financial year their reference number encodes, and now show their release dates in 2020. Their date_precision changes from 2019 to the full date. No entry becomes feed-eligible and no first-seen date moves.

Which reuse pattern breaks. Anyone counting rows per agency or joining on request_id: 19 identifiers stop resolving. Each has one successor, the same request under the agency’s other rendering of its reference number.

2026-09-27 — Distinct entries that shared a reference number are no longer merged (76cc9d5)

Where two entries on an agency’s log shared a reference number, the merge could fold them into one row. This happened when the two were listed on different pages, when one was collected again later, or when an archived copy shared the reference. The merged row could then show one entry’s date with the other entry’s description. PM&C FOI/2014/206 showed a request about the 2013 election with the release date of an unrelated Bureau of Meteorology request. Rows sharing a reference now merge unless their titles name different FOI numbers, or the two were seen on the log at the same time with different descriptions or release dates. Three entries were separated by the collection on 27 September, which saw both entries of each pair on the log together: Defence 022/26/27 and 1885/25/26, and DSS FOI LEX 46667.

47 entries added, 1 removed. No identifiers retired, 46 identifiers added. One pair of ATO rows for the same request, FOI-2026-00793, is now merged. By agency: PM&C +15, Health +7, Defence +6, ASIC +4, AGD +2, Home Affairs +2, Infrastructure +2, NDIA +2, APSC, DEWR, DSS, DTA, Education, NDIS Commission and NHMRC +1 each, ATO −1. This change ships in the same build as the entry “36 entries published twice are merged”. That change removes one DEWR, one DSS and one NHMRC row, so those three agencies’ counts are unchanged.

Each existing identifier stays on one of its entries: the one collected first, or, for 19 identifiers, the one a review matched to the row published under it before. Each separated entry gets its own identifier: the existing one followed by its release date, for example Department of the Prime Minister and Cabinet:FOI/2020/165:2021-09-05. Where two separated entries share a release date, a six-character code follows the date.

43 descriptions change on entries that were not separated. Where the agency edited an entry, the row now shows the version most recently seen on the log, not the longest. Most changes are small: 18 Defence entries lose a trailing “For access to the document”, and others gain or lose a full stop. A few are corrections by the agency: two eSafety entries, LOG-200 and LOG-201, now carry the titles the agency gave them the day after first listing them.

What a consumer would observe. In foiforest.csv, 46 more rows, 43 changed descriptions, and 46 new request_id values. 22 separated identifiers now show a different release date or a different description from the row published under them before, because that row mixed two entries. DSS FOI LEX 46667 shows a different date and description: it again names the 2023 entry it named until 24 September, and the September 2026 decision merged into it on 25 September now has its own identifier. Three existing identifiers now carry a later date_first_scraped: APSC LEX 1655, Defence 938/25/26 and NDIS Commission 195 – 25/26 – (4). In each, the entry that keeps the identifier was first collected later than the entry separated from it. No entry is added to the Atom feeds. The per-agency raw CSVs are unchanged.

Which reuse pattern breaks. Anyone who stored a merged row’s date and description as a pair: in 22 of the 43 merged rows, the two came from different entries. Anyone counting rows per request_id: a request that was one row can now be two, each with its own identifier.

2026-09-24 — 18 duplicate entries merged where an agency re-rendered its reference number (bcc6253)

Health and NDIA both re-render their own reference numbers between collections. Health published one request as FOI 26-3260 on 11 September and FOI 26-3260 -2026 on 15 September; NDIA published another as FOI 25.26-0657 and later FOI 25/26-0657. Reference matching already folded separator differences, but not a trailing year, and the rule that merges an entry the agency has edited compared the reference as raw text. Each pair was therefore published as two entries. Both now match on the normalised reference.

18 entries removed, 18 identifiers retired, none added. Health falls from 1520 rows to 1505, NDIA from 783 to 780. No other agency is affected. Every retired identifier maps to a surviving entry; the mapping is in retired_ids.csv. The retired Health references are 25-0105-2025, and FOI 26-1951-2025, 26-1952-2025, 26-1977-2025, 26-2002-2025, 26-2066-2025, 26-2068-2025, 26-2086-2025, 26-2151-2025, 26-2187-2026, 26-2575-2026, 26-2612-2026, 26-2718-2026, 26-3159 -2026 and 26-3260 -2026. The NDIA references are 25/26-0321, FOI 25.26-0657 and FOI 25.26-0905.

Fourteen surviving entries gain a document link, because the agency renamed the PDF when it retitled the entry and the two copies named different files. Six take the longer of their two descriptions. One entry loses a link: Health FOI 26-1882-2025 falls from two documents to one, and that document is on FOI 26-1882, a separately archived copy of the same request that this change does not merge. No document was dropped from the dataset.

What a consumer would observe. In foiforest.csv, 18 fewer rows, 15 rows whose document_urls count changed, and 6 whose description changed. The per-agency raw/health.csv and raw/ndia.csv files are unchanged: they carry every collected content state, including both spellings of the reference.

Which reuse pattern breaks. An identifier built from a Health or NDIA reference carrying a trailing year, or from an NDIA reference punctuated with a full stop, will no longer resolve. Looking the reference up with the year stripped, or with the separator normalised, finds the surviving entry. Matching by description and date is unaffected.

2026-09-11 — 31 entries stop listing the same document under two addresses (57715d9)

When two collected copies of one entry are merged, their document links are combined. Four of the merge rules combined them by comparing the addresses as text, so an archived copy and the agency’s own address for the same file, or two archive captures of it, were both kept. They now compare by the file the address points at and keep one address per file, preferring the archived copy.

31 existing entries’ document_urls changed; 114 surplus addresses were removed. No entry was added or removed, no identifier was retired, no document was dropped from any entry, and no other field moved. Link counts before and after:

  • ACCC: 1008715/2026-2027 (4→3)
  • AFP: 009-2025, 015-2025, 043-2024, 044-2024, 057-2024 (each 2→1)
  • APRA: 2013/01 (8→6)
  • ASIC: 131 (7→5), FOI 265-2025 (3→2)
  • Comcare: SOLEX13088 (3→2)
  • DCCEEW: 81815 (3→2)
  • Defence: 1025/25/26 (3→2), 843/24/25 (2→1), 904/24/25 (2→1)
  • Finance: 25-26/020 IR (110→55)
  • Health: FOI 26-2182 IR-2026 (4→3)
  • IP Australia: 1107 (4→2), 1108 (2→1), 23/7056 (6→3), 23/7450 (2→1), 24/8721 (2→1), 24/9991 (4→2), 25/11156 (3→2), 25/11201 (2→1), 25/11713 (2→1), 26/12115 (4→2), 26/12726 (4→2)
  • NIAA: FOI/2526/031 (3→2)
  • TGA: 25-0094 (28→16), FOI 3659 (12→7), FOI 3784 (20→13)

What a consumer would observe. In foiforest.csv, 31 rows whose document_urls is shorter than in the previous build. The per-agency raw/*.csv files are unchanged.

Which reuse pattern breaks. Counting documents by splitting document_urls overstated these entries before this build. Matching them on the exact document_urls string against an earlier copy will fail; matching by request_id, or by reference number and date, is unaffected.

2026-09-07 — Australian Federal Police: 14 descriptions regain the spaces between paragraphs (5699758)

The AFP reformatted its disclosure log between 4 and 7 September 2026: a “Date to be removed from website” column was dropped from the current tables, and multi-paragraph request summaries are now emitted as adjacent paragraph and list elements with no whitespace between them. The dataset’s default table reader concatenates such elements, so it began returning summaries with words run together (“Request for:The total cost…”). It also emerged that the AFP’s 2023–2025 summaries had been stored that way since they were first collected, for the same reason. The AFP log is now read cell by cell, which renders each paragraph or list item boundary as a space.

14 existing entries’ descriptions changed. No entry was added or removed by the change, no identifier was retired, and no date or document link moved. Thirteen of the fourteen differ only in whitespace: a space now separates sentences or list items that were previously fused. The fourteenth, 159-2023, additionally loses the literal “-” markers the agency had used before it converted that summary to a list; the new text is what the agency’s page now shows. The affected references: 009-2025, 015-2025, 023-2024, 026-2024, 043-2024, 044-2024, 057-2024, 101-2023, 103-2023, 120-2023, 121-2023, 144-2023, 159-2023, 32-2020.

The remaining 66 of the 80 summaries that the new reader renders differently are unchanged in foiforest.csv, either because the change was invisible after the dataset’s own whitespace normalisation or because an archived copy of the same entry already carried the spaced text and had been preferred.

What a consumer would observe. In foiforest.csv, 14 AFP rows whose description differs from the previous build by inserted spaces (one also by removed dashes); everything else on those rows is identical. In raw/afp.csv, which carries every content state a row has been collected in, 80 rows appear a second time with the spaced description and a scraped_at of 7 September 2026, alongside their original. Ten new entries (035-2026 to 044-2026) arrive in the same build through ordinary collection; they are not part of this change.

Which reuse pattern breaks. Matching AFP entries by exact description text against a copy taken before 7 September 2026 will miss these 14 rows. Matching by request_id, or by reference number and date, is unaffected.

2026-09-03 — 35 entries recovered from archived captures (d10f07c)

Archived captures of 17 agencies’ logs were re-selected page by page. The earlier selection had kept only one capture per crawl day across all of a log’s pages, so the later pages of a paginated log, and older year pages crawled on the same day as the current one, were never read. 35 entries not held anywhere in the dataset were recovered: Health, Disability and Ageing 26 (from the 2010 to 2019 year pages of its legacy log), Infrastructure, Transport, Regional Development, Communications, Sport and the Arts 6, Office of the Australian Information Commissioner 2, Defence 1.

No existing entry was altered and no identifier was retired. The same re-selection also found archived copies of roughly 380 entries the dataset already held from the agencies’ live logs; those were deliberately not added, because merging them would have replaced document links on existing entries (see the entry below for the merge rule that has since changed).

What a consumer would observe. 35 new rows whose date_first_scraped is the day they were added, not the date of the capture they came from; their date fields are the agencies’ own, from 2010 onward. Fourteen further archived rows were identified as the same requests as existing entries under a differently rendered reference or description and were held out; they are recorded in the project and may be added later if the match is confirmed.

2026-08-28 — Australian Electoral Commission: three empty entries removed (9729ccb)

The AEC publishes each release as a heading followed by a description and a list of documents. Three headings on its 2017 and 2021 pages have nothing underneath them at all, and this project was reading each one as an entry — producing a record that carried a reference number and nothing else: no description, no documents, no date beyond the year of the page it sat on.

They were never usable. They are now no longer collected, and the three that had already been published have been removed.

3 entries removed, all from the Australian Electoral Commission: LEX369 and LEX452 (2021), and LS5796 (2017). The agency’s entry count falls from 165 to 162. No other entry changed, in any agency.

3 identifiers retired, one per removed entry, listed in retired_ids.csv with resolution removed:

Australian Electoral Commission:LEX369
Australian Electoral Commission:LEX452
Australian Electoral Commission:LS5796

They have no successor. The entries did not move, merge or get renamed — there was no underlying release to point to.

Corrected 29 September 2026: until then, retired_ids.csv listed these three identifiers with resolution unmatched, not removed. The file now shows removed.

Which reuse patterns break: anything holding one of those three identifiers will stop resolving. Anything counting AEC entries will see three fewer. If you have relied on this project’s AEC data as a complete list of the references appearing on the agency’s pages, note that these three references do still appear there as headings — what is gone is the empty record this project built from them, not anything the AEC has withdrawn.

A note on the 2017 page, which is what surfaced this: it carries the LS5796 heading twice, so this project generated two identical empty records for it. Both are removed.

2026-08-17 — date_first_scraped fixed to Australian time (80ae94e)

date_first_scraped records the date this project first collected an entry. It was derived from a stored timestamp, read in whatever timezone that timestamp happened to be labelled with — and that label belongs to the file, written by whichever machine last saved it, not to the entry. So an ordinary collection run, adding one new entry to an agency’s file, could re-label years of already collected entries and move all of their dates by a day. On 17 August 2026 a routine collection from the Reserve Bank did exactly that to an entry published four days earlier, and thirty-nine agency files were queued to follow.

The date is now always read in Australian Eastern time and the stored label is ignored, so it can no longer move except by re-collection.

888 entries have date_first_scraped one day later than before, across ten agencies: Home Affairs (446), Civil Aviation Safety Authority (207), Comcare (141), Australian Bureau of Statistics (65), Administrative Review Tribunal (14), Commonwealth Ombudsman (4), National Health and Medical Research Council (4), the ACCC and the NDIS Quality and Safeguards Commission (3 each), and the Australian Maritime Safety Authority (1). Every one moved in the same direction, by exactly one day. No other entry’s date_first_scraped changed.

No entries were added or removed, and no identifiers were retired. date_first_scraped is not part of any identifier, and every other field — including each entry’s own dates — is unchanged for every entry, in every agency.

Which reuse patterns break: anything keyed on date_first_scraped, or on the <updated> timestamp of an Atom feed entry, which is derived from it. Those 888 entries now read one day later than in a copy taken before 17 August 2026; the moment of collection itself has not changed. Feed subscribers will see the affected entries re-surface as unread once — the same entries, not new ones.

What is not recorded: how long each entry had been reading a day early. That depended on when each agency file was last written and by which machine, and is not recoverable from the files.

2026-08-14 — Digital Transformation Agency: “Released” restored to descriptions (467b8e8)

The DTA does not publish a summary of each request. What it publishes is a list of the documents released, and each item in that list is titled “Released documents 032”, “Released Document 1”, and so on. Those titles are what this project shows as the entry’s description, for want of anything else.

A rule in the DTA reader removed a leading “Released” from the description. It was meant to strip boilerplate, but here it was taking the first word off a document’s title: entry 032 read documents 032, entry 011 read documents 011, and entry 022 read Document Redacted. Where an entry released several documents the rule removed only the first “Released” and left the rest, so entry 018 read Document 1 Released Document 4 Redacted Released Document 3 Released Document 2. The rule has been removed.

10 descriptions changed, all Digital Transformation Agency, each gaining a leading “Released”. Entries scraped from the DTA’s archived log are unaffected — those carry real request summaries.

No entries were added or removed (29,210 rows before and after) and no identifiers were retired: every DTA entry carries the agency’s own reference number, so its identifier does not depend on the description. No other field changed for any entry, in any agency.

Which reuse patterns break: matching these 10 entries on description text. A copy taken before 14 August 2026 has the truncated wording.

What is not recorded: how long the truncation was published. The rule predates 21 April 2026, when the file it lives in was split into its current form; earlier history was not examined.

2026-08-14 — File sizes removed from descriptions (484ae74)

Agencies routinely append a download’s size to the link text on their disclosure log — “Indian Surrogacy Case (4.5MB PDF)”, “Annual Report (PDF: 2530 KB)”, “Board papers (.zip, 8.9 MB)” — and the whole of that link text was being collected as the entry’s description. The size describes the file, not the request, and it changes whenever an agency re-uploads a document. It is now removed. Only brackets whose entire contents are a size and/or a file format are affected, so a description that mentions a size in its own text (“documents about the 300 MB storage limit”) keeps it.

277 descriptions changed across 15 agencies — Infrastructure (83), Services Australia (77), the AFP (46), Comcare (29) and Home Affairs (12), with ten further agencies at 7 or fewer: the AEC (7), Health (5), the ART (4), the AIHW, CASA and the TGA (3 each), the ATO (2), and AUSTRAC, CSIRO and the DTA (1 each). 291 distinct annotation forms were removed. No description was emptied by the change.

No entries were added or removed — 29,201 rows before and after. Nothing merged: no pair of entries became identical once the sizes were gone.

78 identifiers retired and 78 are new, a one-for-one replacement. All 78 belong to entries with no reference number of their own, whose identifier is derived from the description and is therefore re-minted when it changes — 71 Services Australia, 6 Department of Human Services, and one AUSTRAC. Every one is mapped to its successor in retired_ids.csv. Entries that carry a reference number keep their identifier.

No other published field moved. Every column was compared before and after across the 28,837 identifiers that identify exactly one entry on both sides: dates, date precision, access outcome, document links, reference numbers, agency, notes, contact and collection date are all unchanged.

Which reuse patterns break: matching these entries on description text. A copy taken before 14 August 2026 carries the size annotation where the dataset now does not, so exact-text matching against the 277 entries above will miss. Matching on reference number is unaffected, as is matching on identifier except for the 78 listed in retired_ids.csv.

What is not recorded: how many descriptions carried a size annotation in builds before 14 August 2026. The 277 figure is measured against the build of 13 August 2026; earlier builds were not re-examined, and agencies have edited their own link text over time.

2026-08-11 — Bureau of Meteorology: duplicate copies of two entries removed (2a68e48)

The Bureau of Meteorology reworded several of its disclosure log entries, changing “Bureau” to “bureau” in descriptions and in the notes about what was withheld. The project records an entry’s content as it finds it, so the reworded entries were collected as fresh copies alongside the ones already held, and the 11 August 2026 build published both.

The step that merges such a pair recognises the two copies by checking that they point at the same documents. It compared each copy’s whole document-link field as a single piece of text, which only ever matched when an entry had exactly one document. Entries with several documents were never recognised, so their duplicate copies survived. Five of the Bureau’s reworded entries had one document each and merged as intended; the three with more did not.

Two entries lose a duplicate copy: FOI30/165, dated 25 September 2025, and FOI30/202, dated 24 March 2026. Each appeared twice in the 11 August 2026 build and appears once from now on. The surviving copy carries the Bureau’s current wording and the earlier collection date.

One entry’s document list is halved: FOI30/206, dated 4 March 2026, listed each of its five documents twice — once as a web-archive copy and once as a direct address to the Bureau’s site. It now lists each once, as the archive copy. The five documents are the same five; none is lost.

Two rows are removed and no identifiers were retired. Both removed rows were duplicate copies of entries that remain, and the two copies of each already shared one identifier, so the identifier still resolves to the surviving copy. date_first_scraped keeps the earlier of the two dates, so neither entry reappears as newly added.

Which reuse patterns break: If you counted rows per agency, the Bureau of Meteorology has two fewer. If you counted document links per entry, FOI30/206 has five fewer while releasing the same documents. If you matched these entries on description text, note that the wording changed at the Bureau’s end, not ours: copies taken before 11 August 2026 read “Bureau” where the dataset now reads “bureau”.

What is not recorded: How many multi-document entries elsewhere in the dataset kept a duplicate copy for this reason before 11 August 2026. Across the dataset as it now stands these two are the only ones, but earlier builds were not re-examined.

2026-08-10 — ASIC: an unusable link removed from one entry (2184ef8)

One ASIC entry carried a document link that was not a link. Instead of a web address, the field held an unrendered placeholder from the agency’s content-management system:

https://www.asic.gov.au/umbraco/

/{localLink:umb:/media/adb6c005044143fc95d4755826c2d847}

It contains stray markup and braces, and no browser or tool could ever have opened it. It has been removed.

One entry loses one link: ASIC FOI 202-2025. No other link on that entry changes, and the entry itself is unaffected. No entries were added or removed, and no identifiers were retired.

This is the only occurrence anywhere in the dataset.

Which reuse patterns break: If you counted document links per entry, ASIC FOI 202-2025 has one fewer. If you attempted to fetch every link and recorded failures, this one will no longer appear as a failure — it will not appear at all.

2026-08-10 — Defence: unusable source URLs repaired on 47 entries (13fb70b)

page_url records where an entry was read from. For 47 Defence entries covering 2011-12, that address had been recorded through a web archive under a corrupted host name, produced by the archive’s own crawler rather than by Defence. The address could not be opened.

It now points at the same archived page under its correct address, which resolves. Only page_url and the derived canonical_url change. The entries themselves were confirmed to be genuine records of the 2011-12 Defence disclosure log before the addresses were touched — no entry was added, removed or altered, and no identifier was retired.

Which reuse patterns break: If you resolved page_url for these entries and recorded a failure, that failure will not reproduce.

2026-08-10 — National Archives: missing spaces around a request’s headings restored (b370f0c)

The National Archives of Australia writes some disclosure log entries as a numbered list, each item beginning with a bold heading and continuing into prose. Where such an item also contained a nested list of sub-requests, the space after the heading was being lost during scraping, so the heading ran into the sentence that followed it:

System-Generated Statistical ReportsAs identified on Page 37 of the RAM Legislative Changes Functional Requirements Specification (Stage 3)…

The spaces are now preserved, so the heading reads as its own phrase.

One description changed — National Archives of Australia FOI 255, dated 29 July 2026, with two run-together boundaries separated. No words were added or removed.

No entries were added or removed, and no identifiers were retired. The entry carries a reference number, and request_id is built from the reference number rather than from the description, so it was unaffected.

date_first_scraped is unchanged, so the entry does not reappear as newly added.

Only entries currently on the National Archives disclosure log were re-read. Entries that have since rotated off it were not revisited, so any older National Archives entry written in the same style keeps its run-together text. How many such entries exist is not recorded.

If you matched this entry on description text, matches made before 10 August 2026 may not reproduce. A phrase search spanning one of these boundaries would previously have failed and will now succeed.

2026-08-10 — ATO descriptions: missing spaces between sentences restored (189791c)

The Australian Taxation Office writes each disclosure log description as several lines — what was asked for, then what was released, then what was withheld. The line breaks were being lost during scraping, so the lines ran together with no space between them:

…to deliver messages to FOI/administrative access requestors.Decision letter and documentInformation not relevant to the request excluded.

The breaks are now read as spaces, so the same text reads as separate sentences.

350 descriptions changed, all Australian Taxation Office — 53% of its 656 entries, with 883 run-together boundaries separated. The affected entries are spread across the whole log, from May 2011 to July 2026, with 2017 the heaviest year at 91. No words were added or removed from any description.

No entries were added or removed, and no identifiers were retired. Every ATO entry carries a reference number, and request_id is built from the reference number rather than from the description, so it was unaffected.

date_first_scraped is also unchanged, so these entries do not reappear as newly added.

If you matched ATO entries on description text, matches made before 10 August 2026 may not reproduce. A phrase search that spans one of these boundaries would previously have failed and will now succeed.

2026-08-09 — Repeated text removed from descriptions (3be0742)

Several agencies publish a short title alongside a longer summary, and both were being pasted into description. Where the two said the same thing, the text appeared twice — CASA’s Basair AOC 2006 - 2009 Basair AOC 2006 - 2009. is typical. Text is now dropped when its words already appear, in the same order, elsewhere in the same description, so nothing removed is missing from the entry.

1,015 descriptions changed across 16 agencies — IP Australia (237), Health (225), DCCEEW (161), Education (130), DEWR (89), AMSA (32), Treasury (25), the ART (24) and CASA (23), with seven further agencies under 21 each.

Two entries were removed, 29,158 to 29,156. Both were duplicate pairs that had survived only because their two copies were worded differently; with the repeated text gone the wording matched and they collapsed. One is a Health entry whose archived copy merged into its live copy.

21 identifiers retired and 19 are new. The counts differ because two of the retired identifiers belong to the entries removed above and have no successor. The other 19 — 18 CASA, one Education — are entries with no reference number, whose identifier is derived from the description and was re-minted when it changed.

Corrected 29 September 2026: the two retired identifiers of the collapsed entries do have a successor, the surviving copy of each entry. DEWR nref:62bbd1dac3e1 maps to LEX 1160, 1162, and Health 25-0009 LD maps to 25-0009 LD-2024. Until then, retired_ids.csv listed both as unmatched. It now lists both as mapped, with those successors.

Seven DEWR entries kept a different copy of themselves. Each was already a collapsed duplicate pair, and the shorter description changed which of the two copies survives: six changed page_url and notes, and one changed date_first_scraped from 20 July 2026 to 25 April 2026. The request itself — description, date, reference number, documents — is the same either way.

If you matched entries on description text, matches made before 9 August 2026 against these agencies may not reproduce.

2026-08-05 — date precision published as data (57fc541)

Schema addition: two new columns, date_precision and date_basis, are now published in foiforest.csv, appended after subpage_url so a parser reading columns by position is unaffected. A parser asserting an exact column count will see 16 rather than 14.

What the fields carry: date_precision records how precisely the source stated the date behind date, as a truncated ISO string — 2016, 2016-09 or 2016-09-01; date_basis records whether those components are a calendar year or an Australian financial year.

What was not affected: no existing value changed — no date, date_source, description, agency or request_id — and the row count is unchanged at 29,100, with no identifiers retired and none new.

A caution specific to financial_year: the Fair Work Commission assigns a reference number on receipt, not on decision, so for those 92 entries the year describes receipt and the release followed at an unknown interval afterwards. An entry stating only “2010” has a release date somewhere within that year; an FWC entry carries no release-date evidence at all.

Who this matters to: if you were using date_source to judge how precise a date is, date_precision is the field that answers that.

Related, not part of this entry: datapackage.json, now published beside the CSV at https://data.foiforest.org/datapackage.json, carries the full field descriptions in machine-readable form.

2026-08-03 — APRA date precision labels corrected

Correction: 100 of APRA’s 148 entries were labelled as having month-only precision when the agency had in fact published a full date. Those 100 entries carried date_source = "date_published_month" and date_precision: month only in the notes field, both wrong.

APRA’s disclosure log mixes two date formats in the same column — some entries give a full date (“9 May 2016”), others only a month (“January 2026”). The code that flags the month-only entries flagged every entry instead, without checking what the agency had actually published.

What was affected: the date_source field and the date_precision: key inside notes, for 100 entries at one agency — about 0.3% of the dataset’s 29,075 entries as at 3 August 2026.

What was not affected: no date value changed. date, date_of_access and date_published are the same before and after, and every entry keeps the same request_id. The dates themselves were always correct. Only the labels describing how precise they were were wrong.

Who this matters to: if you excluded date_published_month entries as too imprecise to use — a reasonable thing to do — you silently dropped 100 APRA entries that had full dates and should have been kept. If you used date directly, nothing you did was affected.

How long it was wrong: the labelling was introduced on 6 April 2026 (9ea5840), hours before the first public data files were generated the same day (ed91364). It was therefore present in every published version of the dataset until this correction, a period of about four months.

Identifying the affected entries in a copy you already hold: before the fix, all 147 APRA entries with a date carried date_source = "date_published_month". 100 of those were wrong and 47 were right, so the label alone does not identify them — filtering on it and relabelling everything would introduce the opposite error in 47 entries.

  • 94 of the 100 can be identified: date_source = "date_published_month" and a date whose day is not the 1st. A month-only date is always floored to the 1st, so any other day proves the agency published a full date.
  • 6 cannot be identified: they are genuinely dated the 1st — 2012-06-01, 2015-12-01, 2019-07-01 (three entries) and 2021-09-01 — and in a published copy a stated 1st and a floored 1st are identical. All 47 correctly-labelled month-only entries are also dated the 1st, so there is nothing to separate them by.

Those six are the reason the fix reads the raw date string the agency published rather than inferring precision from the parsed date. A rule based on the parsed value would have corrected 94 and left 6 wrong.

If you hold an old copy, re-download it rather than patching it.


Entries below this line are reconstructed

Everything above was written at the time the change was made. Everything below was reconstructed on 3 August 2026 from the project’s repository history and by comparing retained internal builds. It was not recorded contemporaneously, and it should be read as a reconstruction rather than a record.

Two limits on accuracy. First, changes are attributed by comparing the last build before a change with the first build after it, so where several changes landed close together the effects cannot always be separated — this is noted where it applies. Second, those windows sometimes contain an ordinary daily scrape, so a small number of the changes counted below are normal re-scraping rather than the code change. Counts are therefore upper bounds unless stated otherwise.

Where something could not be established from the record, it says so rather than estimating.

Row counts and identifier counts move independently, and the numbers below will not subtract to each other. Because identifiers are derived from an entry’s content, a change can retire an identifier without removing an entry — the entry survives under a new identifier — and it can remove entries without retiring identifiers, when a row collapses into another that already shared its identifier. Both happened.

The clearest case is 25 April, where the arithmetic is exact. 231 rows sat on identifiers that were retired, 128 rows arrived on new identifiers, and rows carrying identifiers that survived the change fell by 180 on their own — because reassigning records between departments gave them the same identifier as an existing entry, which then collapsed. Those three movements sum to the −283 change in the total.

Note also that a count of identifiers and a count of rows are not the same measurement: on 25 April, 230 identifiers were retired but 231 rows carried them, because one identifier covered two rows.

Where an entry’s identifier changed rather than the entry being removed, the mapping is in retired_ids.csv, described above.

2026-07-21 — Duplicates collapsed where an agency re-rendered its own references

Commit 52059b8. Build of 21 July, afternoon.

Agencies change how they write their own reference numbers over time, so the same request could appear twice — once as 17/138 from an older capture and once as FOI 17/138 from a newer one. These pairs were collapsed into single entries.

43 entries removed, 43 request_ids retired, none added. Affected: Finance (29), Employment and Workplace Relations (7), Health (3), Services Australia (3), Defence (1).

Dataset total: 29,223 → 29,180 entries.

Retired identifiers: 43, all still absent today and listed in retired_ids.csv under 2026-07-21-reference-variants. 40 are mapped to the entry that survived the collapse; 3 could not be matched automatically.

This was a tightly bracketed change — the only commit between the two builds compared — so the attribution here is firmer than for most entries below.

2026-07-21 — Services Australia duplicates collapsed; APRA reader rebuilt

Commits ede69cc (duplicate handling), b22b989 (APRA). The two commits landed on 20 and 21 July between the same pair of builds and cannot be separated.

Net 78 entries removed. 148 request_ids retired, 75 new ones appeared. 145 of the retirements are Services Australia, where entries captured under the department’s former name (Department of Human Services:...) were matched to their equivalents and collapsed. Document links changed for 130 entries, mostly Employment and Workplace Relations (113).

Dataset total: 29,295 → 29,217 entries.

Retired identifiers: 148, all still absent today and listed in retired_ids.csv under 2026-07-21-services-australia. 75 are mapped to the entry that survived the collapse; 73 could not be matched automatically.

APRA’s disclosure log was redesigned by the agency and the reader was rewritten to match. The previous reader had been returning nothing since some point after 6 May 2026 — the exact date the agency changed the page is not recorded — so APRA entries published in that period were not being updated. No previously published APRA entry was altered by the rewrite.

2026-07-19 — NDIA descriptions restructured; description cleaning fixed

Commits c87b1c5 (NDIA), b3d18fb (cleaning).

NDIA release notes had been appended to the description field. They were moved into notes, where the rest of that kind of material lives. 752 NDIA descriptions changed, along with 753 notes values. No entries were added or removed and no identifiers changed.

Separately, the routine that strips boilerplate (“please email … to obtain documents”) from descriptions had been truncating some descriptions at the wrong point. 287 descriptions changed across Home Affairs (71), the Federal Court (50), the Human Rights Commission (27), Defence (21), the Ombudsman (15) and others. Three Infrastructure entries changed identifier as a result, because identifiers for entries without a reference number are derived from the description.

If you matched entries on description text, matches made before 19 July 2026 against these agencies may not reproduce.

2026-06-09 to 2026-07-15 — Smaller duplicate-handling changes

Commits c05b1af, 76390cb, 9fcb63f, e751da9. Several rounds of adjustment to duplicate detection and to Infrastructure reference numbers. Their combined effect on already-published entries was small: one Australian Research Council identifier retired (its content survives under another identifier), document links changed for 8 entries (6 Infrastructure, 2 AFP), and one Federal Court note changed. No other published entry was altered.

2026-05-20 — Revision parameters stripped from URLs

Commit bd80304. A rev= parameter in some agency URLs changed on every scrape, making unchanged entries look like new ones. It is now stripped. Document links changed for 13 entries, all IP Australia. No entries added or removed.

2026-04-25 — Records reassigned between successor departments

Commit 4851576, together with several archive sources added the previous day (be8c2fe, 9eb1744, 0e16a46, 268a92d) — the effects cannot be fully separated.

Some archived disclosure logs carry records belonging to more than one of today’s departments, because the department that published them has since been split. Records were reassigned to the correct successor: 232 entries changed agency, moving between DCCEEW and DAFF, and between Education and Employment and Workplace Relations.

283 entries were removed as the reassignment allowed previously-invisible duplicates to collapse. 230 request_ids were retired and 128 new ones appeared. The identifier is built from the agency name, so an entry that moved department kept its content but changed its identifier.

Dataset total: 25,974 → 25,691 entries.

If you were tracking entries at these departments by request_id, identifiers recorded before 25 April 2026 may no longer resolve even where the entry itself is still present. Of the 230 identifiers retired, 228 are still absent today and are listed in retired_ids.csv under 2026-04-25-successor-departments; 192 of those are mapped to their current identifier, 36 could not be matched automatically. (The other 2 have since reappeared in the data and are not listed.)

2026-04-24 — CASA reference numbers corrected

Commit 5fa5f23, in a window that also contains a change of date basis (02051be) and several new archive sources — the effects cannot be fully separated.

CASA reference numbers had been captured with the request description appended to them (for example, F25/24242 Records of UAP UFOs NHI as a single reference number). Every affected reference number was corrected.

All 186 CASA request_ids in the dataset were retired and replaced, since the identifier is built from the reference number. No CASA entries were lost; they are all present under corrected identifiers. Listed in retired_ids.csv under 2026-04-24-casa-references: 179 of the 186 are mapped to their current identifier, and 7 could not be matched automatically because the entry’s description also changed.

226 new identifiers appeared against the 186 retired. That is the 186 replacements plus 40 belonging to CASA entries that arrived from an archived CASA source added in the same window (8b41a49), which is why the dataset total rises by 40 in a change that removed nothing. The two cannot be fully separated.

In the same window, document links changed for 278 Treasury entries and notes for 6, from archive sources added at the same time.

Dataset total: 24,315 → 24,355 entries.

2026-04-23 — Agency renames applied before duplicate detection

Commit 55e4044. Machinery-of-government renames were moved earlier in processing so that records from a department’s former name could be matched against its current name. Comparing the builds either side shows no entries added or removed and no identifiers changed; document links changed for 42 Industry entries.

2026-04-06 to 2026-07-15 — Agencies and archive sources added

Coverage grew substantially over this period. New agencies added after publication began: Environment (22 April), Education, Industry/RET archives (23 April), FaHCSIA (24 April), CSIRO, Federal Court, RBA (28 April), ACIC (1 May), AEC, TGA, Bureau of Meteorology (3 May), AIHW, APSC, ACQSC (14 May), Jobs and Skills Australia, NHMRC, ARC, Fair Work Commission, Australia Post, NBN Co, CDPP (8 June), National Archives (8 July), Communications archive (15 July). Archived sources were also added for a number of agencies already covered, including Treasury, DEWR, DTA, Comcare, CASA, the MRT/RRT, the ART, Customs (ACBPS), the Fair Work Ombudsman, IP Australia, Infrastructure/DCA, Defence, DVA and Home Affairs.

These are additions, and are listed only because adding an archived source for an agency already covered can change existing entries — the added source may supply a document link or a reference number that the live log did not, and can cause an existing entry to be merged with a newly-arriving one. The individual effects are not separately recorded.

2026-04-06 — First publication

Commit ed91364. The public data files were generated for the first time. Anything before this date was not published.

Changes with no observable effect on published entries

Commits 9c9c8de (5 May, a Bureau of Meteorology duplicate correction) and a29a1e6 (8 June, duplicate-handling logic) were compared across the builds either side and changed nothing in the published files — no entries removed, no identifiers retired, no field values altered. What the BOM correction was intended to fix is not recorded.


Changes affecting published rows have been recorded contemporaneously since 3 August 2026. Corrections and questions: see the About page.