Download the FOI Forest dataset

The searchable, centralised archive of Australian FOI releases

FOI Forest collects disclosure log data from — federal agencies and processes it into a single dataset with a consistent structure. The search on this site runs on that dataset, and you can download it here. If you prefer to work with the data as scraped, before any processing, separate files for each agency are also available.

All datasets are released under CC0 1.0. You can use them for any purpose without restriction. Attribution is appreciated but not required. CC0 applies to the dataset only. The documents it links to, including archived copies, remain under each agency’s copyright terms.

Unified dataset

↓ Download foiforest.csv

Data last rebuilt 6 October 2026, 07:12 UTC.

The timestamp above is the time the current build started. The CSV’s HTTP Last-Modified header holds the time the file was uploaded, a few minutes later. If you need to cite or reproduce your results, keep your own copy and record the Last-Modified and ETag headers from your download. Two downloads with the same ETag are the same file. The command curl -sI https://data.foiforest.org/foiforest.csv prints both headers.

The unified dataset contains — records from — agencies, normalised into a consistent schema. Every hour from 8am to 6pm Melbourne time, Monday to Friday, each agency’s disclosure log is checked automatically for new entries. When new entries are found, the dataset is rebuilt from all the data scraped so far, and the published files are replaced. Every Wednesday morning, every page of every live log is also scraped in full, and the files are rebuilt after that check too. The files can therefore change several times in one day. I keep earlier builds privately but do not publish them. I record changes to entries that were already published in the data changelog.

Schema

Field Description
request_id Identifier for the record, unique per row since 28 September 2026. For a record with a reference number it takes the form {agency}:{reference_number}, where {agency} is the agency name recorded for the log the record came from. That name can be an earlier name or a predecessor’s name, so it does not always match the agency field. For a record without a reference number it takes the form {agency}:nref:{hash}. The hash is the first 12 characters of a SHA-1 of that name, the date and the description. An undated record uses its date_first_scraped instead of a date. Where several rows share a reference number, one row keeps the plain identifier. Each other row adds its release date, for example Department of Veterans' Affairs:77505:2026-09-05, and a six-character code where two rows share that date. The identifier stays the same from build to build. A change to the reference number, or to the date or description of a record without one, gives the record a new identifier. The data changelog explains how to find the new identifier.
reference_number The FOI reference number exactly as published by the agency, including its punctuation and any prefix. Empty for records from agencies that do not assign one (notably AUSTRAC). Not unique within an agency.
date The best available date for the request: date_of_access where present, otherwise date_published. Where only partial date information was published this is floored to the earliest plausible day; see date_precision and date_basis. Present for nearly all records.
date_source Where date came from. It records the source, not the precision. date_of_access: taken from the agency’s decision/release date column; date_published: the date the entry appeared on the disclosure log; date_published_month: a month-level publication date; ref_year: a year extracted from the reference number; url_path: a year or month from a dated path in a document link or in the address of the log page, or, for a few TGA entries, the most common year of the other entries on the same log page; year_of_access: a year taken from a heading in the page HTML; none: no date recoverable.
date_of_access The date the applicant received the decision or documents, as recorded by the agency. Empty where the agency does not publish this field.
date_published The date the entry appeared on the agency’s disclosure log. Empty where the agency does not publish this field. See date_first_scraped for an approximate substitute.
date_first_scraped The date, in Melbourne time, when FOI Forest first collected this record. For records that appeared after an agency was first scraped, this approximates the publication date, usually to within a day. It can be later than the true date: logs are checked only on weekdays, some agencies’ older pages are read only in the Wednesday check, and collection stops while a scraper is broken. Records collected in an agency’s first scrape all share that scrape’s date and say nothing about when they were published.
description The request description as published by the agency.
access_outcome Standardised access outcome: Full (full release), Partial (partial release or some documents withheld), or empty. Many agencies do not publish outcome information; empty means not stated, not refused.
document_urls Links to released documents, where published. Multiple links are joined with a semicolon padded by a space on each side. A link can itself contain a semicolon, so split on a semicolon that has whitespace on at least one side, not on a bare semicolon, and trim each part. Some links point to a web page about the release rather than to a document. From time to time FOI Forest replaces an agency’s link with an archived copy of the document, and not for every agency. An archived link starts with https://web.archive.org.au/awa/ (Australian Web Archive) or https://web.archive.org/web/ (Wayback Machine), followed by a 14-digit capture time and the original address. The agency’s original link is kept in the document_urls_live column of the per-agency raw files. Archived links that could not be retrieved are removed. Where an entry was collected more than once, the links from every copy are combined, so one document can appear under two addresses. Links that used http are rewritten to https. Links may no longer resolve, and some agencies supply documents on request rather than linking them.
notes Agency-specific fields that do not map to the standard schema but are worth preserving for search and analysis, as key: value pairs separated by a vertical bar padded with a space on each side.
agency Full agency name, always the agency’s current one, applied uniformly to historical records including those predating that name.
canonical_url URL of the agency disclosure log page this record was found on, with pagination parameters stripped for stability.
subpage_url URL of the individual entry page, where the agency publishes records on separate pages rather than in a single table. Empty otherwise.
date_precision How precisely the source stated the date, as a truncated ISO string: 2016 (year only), 2016-09 (month only), 2016-09-01 (full date). Records what was stated, not a guaranteed interval containing the true date. Read with date_basis.
date_basis Which calendar date_precision’s components use: calendar or financial_year. For financial_year, 2016 means FY2016–17 and date is floored to 1 July. Only Fair Work Commission records carry it; their reference numbers encode the financial year a request was received, not when it was released.
presence_status Where the record can be found on the agency’s website, from an automated weekly comparison with the agency’s current site. live: on the disclosure log, or on a page the log links to. unlinked_page: on a page of the agency’s website that its disclosure log does not link to. listed_in_attachment: listed in a downloadable file on the agency’s website, outside its disclosure log. archive_only: seen in web archive copies or an earlier check, and not found on the agency’s current website. It can be wrong where the agency changed the reference number, published the entry on a page FOI Forest does not read, or moved it to another agency’s log. not_measured: the latest check could not confirm it, because a page did not load or has not been checked yet. Check an individual record before relying on its status. See “No longer on disclosure logs” on the Data page.
reference_base reference_number without a part suffix, where an agency lists one request’s documents as several entries: “FOI-2024/0423164157 - Part 1 of 2” and “- Part 2 of 2” both have “FOI-2024/0423164157”. Each part keeps its own request_id. Count distinct values of agency and reference_base to count requests rather than entries. Identical to reference_number for every other record; empty where that is empty.

The same descriptions are published in machine-readable form as a Frictionless Data Package descriptor.

Notes on the data

Date quality. Dates come directly from agency logs, but date formats vary considerably across agencies and have changed over time within agencies. I corrected some dates by hand where the source had a clear typographical error, such as an invalid month name or transposed digits. Where an agency gives only a year or a month, the date is set to the first day of that period. The field date_precision holds how precisely each date was stated: a year (2016), a month (2016-09) or a full date (2016-09-01). For Fair Work Commission records, date_basis shows that a year is a financial year, and the date is set to 1 July. The field date_first_scraped holds the date, in Melbourne time, when an entry was first collected. All entries collected in an agency’s first scrape share that date.

Reference numbers. Most agencies assign a reference number to each FOI request, but not all. Reference numbers are reproduced as published, apart from the removal of extra whitespace and invisible characters, and a few fixes for individual agencies, such as the removal of the “Reference number:” prefix that the Fair Work Commission adds. A small number of agencies appear to have assigned the same reference number to genuinely distinct requests. If you rely on reference numbers to deduplicate or link records, treat them with some caution. Some agencies list one request’s documents as several records, each with a part in its reference number, such as FOI-2024/0423164157 - Part 1 of 2. Each part has its own request_id. The field reference_base holds the reference number without the part, and equals reference_number for every other record. Counting the distinct combinations of agency and reference_base gives an approximate count of requests, because a few agencies have reused a reference number for different requests in different years.

Presence status. The field presence_status shows whether each record was found on the agency’s website at the latest weekly check. It comes from an automated comparison, so it can be wrong for an individual record. No longer on disclosure logs explains the values and their limits.

Access outcomes. Agencies use inconsistent language to describe outcomes, so the standardisation of access_outcome to Full or Partial is approximate. Many records have no outcome, either because the agency does not publish one or because it could not be parsed reliably. An empty value means the agency did not state an outcome. It does not mean the request was refused.

Document links. Agency links to released documents often break. From time to time I replace an agency’s link with an archived copy of the same document from the National Library of Australia’s Australian Web Archive or the Internet Archive’s Wayback Machine. I do this occasionally, and not on every build. I do not replace links for ten agencies whose logs appear to be complete and which, as far as I can tell, do not remove entries or documents: the ACCC, ABS, AUSTRAC, Reserve Bank of Australia, Department of the Prime Minister and Cabinet, Commonwealth Ombudsman, National Indigenous Australians Agency, National Anti-Corruption Commission, Australian Electoral Commission and Department of Home Affairs. Their own links are likely to keep working.

Before a link is replaced, the archive’s index is checked automatically for a capture of the document’s address that the archive recorded as successful. The capture itself is not opened. An archived link can therefore occasionally lead to an error page that the agency’s site returned with a success code. An archived link starts with https://web.archive.org.au/awa/ or https://web.archive.org/web/, followed by a 14-digit capture time and the original address. The agency’s original link is kept in the document_urls_live column of the raw files.

Archived sources. For some agencies, records are also drawn from snapshots in the Australian Web Archive or the Wayback Machine, as well as from the current live log. The snapshots extend coverage back beyond the live log. Document links for these records point to archived copies, which are more likely to stay accessible. The sources for each agency section lists every archived source with the years of its entries.

Coverage limits. The dataset holds only what agencies publish on their disclosure logs. Under section 11C of the Freedom of Information Act 1982, an agency does not have to publish released information that includes personal information or business affairs information where publication would be unreasonable, or information that it is not reasonably practicable to publish. Those releases are absent. Some releases never appear on a log at all. Refused requests are absent, because a disclosure log lists releases only. Agencies not listed under sources for each agency are not covered.

Coverage of an agency can also fall behind. In the agency table on the home page, a note appears on an agency’s row when its log has not been read for more than four days. A QUIET label appears when an agency has gone longer than its own usual interval, and at least 60 days, without a new entry. A quiet agency may simply have published nothing new. A gap can also mean the scraper has stopped working, for example after an agency moves its log or changes its page layout. Before relying on an absence, check the agency’s own log.

Agency names and machinery of government. The agency field in FOI Forest reflects the agency’s current name, applied uniformly across all historical records, including those that predate the current name. I made this choice because agencies handle their own disclosure log history inconsistently. Some carry their full history forward under the current name with no indication of any discontinuity, whereas others start fresh when restructured and drop older records from their log entirely. A small number do indicate where entries come from a predecessor, but even these have inconsistencies. I have no reliable way to determine which historical predecessor a given record belongs to, so using the current agency name seems to me the least bad option.

As a result, filtering by agency name may not return a coherent institutional unit, particularly for departments that have gained or lost functions over time without changing name. The canonical_url field holds the address of the log page where each record was found. If that page is on an archive’s domain or a predecessor’s domain, the record probably predates the current department’s form.

Some agencies have undergone complete institutional succession across multiple names. The Administrative Review Tribunal is the clearest example. I attribute records from the Migration Review Tribunal, the Refugee Review Tribunal and the Administrative Appeals Tribunal to the ART, reflecting a continuous institutional function across three names. Similarly, records from the Department of Human Services appear under Services Australia.

The Department of Infrastructure, Transport, Regional Development, Communications, Sport and the Arts absorbed the former Department of Communications and the Arts in 2020. I attribute records from the communications department’s own disclosure logs, captured between 2011 and 2021, to the current department.

Some agencies are harder to deal with, because two currently distinct agencies share a predecessor that combined the functions of both. This affects the Department of Agriculture, Fisheries and Forestry (DAFF) and the Department of Climate Change, Energy, the Environment and Water (DCCEEW), and also the Department of Education and the Department of Employment and Workplace Relations (DEWR). In these cases the same record often appears in the logs of both agencies after the split, producing duplicates in the dataset.

To reduce this duplication, I attribute a DAFF record to DCCEEW when its reference number also appears in the DCCEEW log, and the DCCEEW record is dated within two years of it or cites its reference number. Where the two copies then match, one is removed during deduplication. For Education and DEWR, a field called creator in the raw scraped data names the predecessor that published each record. Each record is routed by this field to the successor department that took over that function. In the unified dataset the value appears in notes as creator: …. These changes apply only to the unified dataset. The raw files reflect the data as scraped and contain the duplicate records.

I made these attributions only to reduce duplication, and I do not intend them as statements about institutional history. If you are researching an agency that has been through machinery of government changes, also search the agencies that shared its functions at the time.

Raw files

Raw files are available for each agency. They contain the data as scraped from each agency’s disclosure log and from any archived snapshots, before normalisation, date parsing, deduplication or schema standardisation. Column names and structures vary by agency and reflect each agency’s own disclosure log format. Each file covers every source for its agency, including archived logs and predecessors’ logs. Rows that repeat exactly are removed, and the first one collected is kept. Earlier versions of edited entries remain in these files.

The raw files are uploaded in the same build that produces the unified dataset, so the two stay in sync.

Loading…

Processing

Processing has three stages.

Scraping

Each agency’s disclosure log is scraped automatically every hour from 8am to 6pm Melbourne time, Monday to Friday. For some agencies, only the most recent pages of the log or the current year are scraped each hour. In the full check on Wednesday mornings, every page of every live log is scraped, including older pages that the agency no longer updates. Many agencies block simple HTTP requests, so a headless browser is used where needed. Structured data is extracted from tables, list elements or custom page layouts, depending on the agency.

Transformation

Each agency’s raw data is processed through a common set of steps for date parsing, URL normalisation and schema mapping. Agencies use at least seven distinct date formats between them, and many logs mix several. Most dates come from the agency’s decision date, and otherwise from the date the agency published the entry. For a few agencies that give neither, the date comes from the reference number, a dated path in a document link, a year heading on the log page, the address of the log page, or the most common year of the other entries on the same page. The source of each date is recorded in the date_source field. Most agency-specific fields that do not fit the standard schema are folded into the notes field.

Why deduplication is needed

The same FOI request can appear in the scraped data more than once. Agencies edit existing entries: they reword descriptions, reformat reference numbers and add document links. As at 29 September 2026, Health and the TGA edited entries most often. About 3.6% of Health’s entries and 5.5% of the TGA’s have an edited description. The NDIA reformatted many existing entries at once on 8 June 2026. Some agencies also list one request twice, publish documents in tranches, or keep an entry on both a live log and an archived snapshot.

Each time an agency edits an entry, the edited version is stored as a new row. Without deduplication, the edited version would appear as a new entry, and the earlier version would appear to have left the log. During deduplication, the versions are matched and one entry is kept, with the agency’s latest wording. Earlier versions of edited entries remain in the raw files. foiforest.csv holds only the latest version, with no record of the edit.

Deduplication

The rules are complex, and I am still refining them. If you need certainty about a specific record, check the raw files. Within each agency, duplicates are resolved as follows:

  • An archived copy of an entry that is also on the live log merges into the live entry, when the description and date match or the reference number matches.
  • The versions of an entry that the agency has edited merge into one entry, with the agency’s latest description and every document link from each version. Versions are matched by reference number and log page. They are kept apart when they name different FOI numbers, or when they appeared on the log at the same time with different details.
  • Where an agency has changed the format of a reference number, such as its spacing, entries with the same description and date merge.
  • Records with the same reference number, description and document links collapse to one record, with the earliest date. The description match ignores case, spacing and punctuation.
  • Records with the same reference number, description and decision date but different document links merge, with the links combined.
  • Records with the same reference number and description but different dates and different document links stay as separate rows, to preserve the history of releases in tranches.
  • Records stay as separate rows when one record’s description contains another’s reference number, because they are distinct requests.
  • Where an agency lists the same entry twice, with the same date and substantially the same subject, the two merge.
  • A small number of pairs that I have reviewed by hand merge.
  • Everything else with a duplicate reference number stays as separate rows.

Request IDs

The request_id field gives each row its own identifier. It takes one of two forms:

  • {agency}:{reference_number} for a record with a reference number, for example Department of Veterans' Affairs:77505. {agency} is the agency name recorded for the log the record came from. It can be an earlier name or a predecessor’s name, so it does not always match the agency field. For example, Health records start with Department of Health and Aged Care:, and DAFF records attributed to DCCEEW keep their DAFF name.
  • {agency}:nref:{hash} for a record without a reference number. The hash comes from that same recorded name, the date and the description. For an undated record, the date it was first collected is used instead.

Where several rows share a reference number, one row keeps the plain identifier. Each other row adds its release date, for example Department of Veterans' Affairs:77505:2026-09-05, and a six-character code where two rows share that date. Since 28 September 2026 every row has its own request_id, so you can use it as a key.

An identifier stays the same from build to build. A change to the reference number gives the record a new identifier, and so does a change to the date or description of a record without a reference number. A department’s change of name does not change its identifiers, because the prefix keeps the name recorded for the log. Before 28 September 2026, rows that shared a reference number also shared a request_id. The data changelog explains how to find an identifier’s replacement, and how to tell those rows apart in older copies.

No longer on disclosure logs

The home page shows a figure labelled “No longer on disclosure logs”. It counts the records that are no longer on the agency’s disclosure log: those whose presence_status is archive_only, unlinked_page or listed_in_attachment. Some of these records are still elsewhere on the agency’s website: unlinked_page records are on a page the log does not link to, and listed_in_attachment records are listed in a downloadable file. The field presence_status in foiforest.csv gives the status of each record, and the schema defines its values.

For some agencies, the hourly scrapes read only the most recent pages of the log, so an entry missing from an hourly scrape may be on a page that was not read. The weekly check, usually on Wednesday, reads every page of every live log, and the statuses come from the latest accepted check of each log. These checks began on 20 July 2026. If a check is incomplete or implausible, for example because the number of entries dropped sharply, the previous accepted check is used and the home page notes its date. If one page of a log fails to load, the records last seen on that page are not_measured instead. Edits to the wording of an entry with a reference number do not count as removals, and neither do entries first collected on or after the date of the latest check.

The statuses come from an automated comparison, and archive_only can be wrong. A record is archive_only when it was seen in a web archive copy of the log or in an earlier check, and is not found on the agency’s current website. That status is wrong where the agency changed the reference number, published the entry on a page FOI Forest does not read, or moved it to another agency’s log. A check in October 2026 found a possible match on the agency’s website for fewer than 1 in 100 archive_only records. Treating all of them as matches would change no agency’s share by more than half a percentage point. Check an individual record on the agency’s website before relying on its status.

The figure also misses removals. An entry that an agency removed before any web archive captured its log is not in the dataset at all.

A disclosure log does not say why an entry left it, and I cannot tell. An entry may have been moved, corrected or withdrawn. Entries counted as no longer on disclosure logs remain in the dataset and in search. I do not yet publish a list of removals confirmed by hand.

Sources for each agency

Each agency below is listed with the predecessor agencies whose records are included, the live disclosure log pages that are scraped, and each archived source. For each archived source, the years are those of the earliest and latest entries collected from it. Where an agency has moved its log, the former address is listed too. Each archived source links to the archive’s full list of captures. This list is generated automatically from the scraping configuration. The agencies’ own disclosure logs are the authoritative sources for current data. If something looks wrong, check the source log first.

Loading…

Contact

Found an error or want to suggest an agency to add? Reach me at gabrielle@foiforest.org.