Presentation
AI in Research: Methods and Policy
How can researchers use AI in work that others can check? We examine how to keep a clear record of what the AI did and which evidence supports the result. The researcher remains responsible for every decision.
One part of this work turns policy into decisions researchers can act on. Our university guidance sets boundaries for appropriate AI use and explains what researchers need to disclose. Our research on institutional frameworks compares how universities developed guidance that permits useful work while keeping human oversight.
We also build and document methods. Our document-discovery work uses language models to classify archived web pages. Each classification retains a supporting quotation, and the packaged process lets another researcher inspect or repeat it. A related methodological study asks what researchers must design around a model to make AI-assisted work reliable. Together, these projects connect practical tools with clear responsibility for their use.
Preprint
Read abstract
Companion paper B (submitted to Royal Society Open Science, which uses open peer review): a methodological case study of large language model (LLM)-assisted research, arguing that reliability must be engineered into the human–AI system rather than expected of models alone. This component holds the preprint, electronic supplementary material, and the AB+ literature-grounding pipeline and outputs; shared data live in the Software Tools Dataset and LLM Research-Assistance Corpus components.
Software
Read abstract
LLM Document Discovery v0.2.0 — Apptainer Container Pipeline
Reproducible pipeline for classifying historical web documents (1996–2005) using large language models. Extracts linguistic and structural features from children's web
content archived by the Internet Archive, producing a structured SQLite database of classifications with supporting blockquote evidence.
This release adds a containerised execution pipeline using Apptainer/Singularity, enabling reproducible deployment on both local GPUs and HPC clusters (NCI Gadi).
What's new in v0.2.0:
Apptainer container: wrapping vLLM v0.19.0 + llm-discovery for fully offline, reproducible execution on HPC compute nodes
Single-command deployment: deploy assembles the data directory, syncs to HPC, and submits the PBS job
Multi-model support: tested with google/gemma-4-E4B-it (local RTX 4090), google/gemma-4-31B-it (Gadi V100), and openai/gpt-oss-120b (Gadi H200)
Prompt engineering: rationale-first instruction format achieving 98% structured output compliance (up from <1%)
Crash-safe resumability: killed containers resume from the last completed document–category pair
CLI commands: build, init, download-model, deploy, status –watch, retrieve, for end-to-end HPC lifecycle management
Offline tokeniser support: tiktoken vocab files baked into the container for OpenAI gpt-oss model family
Associated paper:
Johnston, V.H., Ballsun-Stanton, B., Jensen, H.S., Kjelsen, C.K., & Thøgersen, J. (2026). Exploring the Archived Web through AI-Assisted Document Discovery. https://github.com/WEB-CHILD/exploring-the-archived-web-through-ai-assisted-document-discovery
Paper
Paper
Guidance
Read abstract
The use of generative AI in research needs to align with the principles and responsibilities outlined in the Australian Code for the Responsible Conduct of Research 2018. Guidelines about the appropriate disclosure and documentation to accompany the use of Generative AI in research also need to recognise the practical considerations of current or innovative research practices. An approach combining these elements is detailed in this document for Macquarie University researchers.
Presentation
Read abstract
Slides for the Research Data Alliance event:
Title: AI in Action: How Researchers Leverage AI (Asia/Oceania Friendly Time)
https://www.rd-alliance.org/event/ai-in-action-how-researchers-leverage-ai-europe-americas-friendly-time-2/
Presentation
Presentation
Read abstract
A briefing on the Macquarie University Generative AI in Research Guidance Note. The Guidance note can be found at https://policies.mq.edu.au/download.php?associated=1&id=768&version=1
Guidance
Read abstract
This is the university policy/guidance note for the use of Generative AI and LLMs at Macquarie University, Australia. The original source is https://policies.mq.edu.au/download.php?associated=1&id=768&version=1
Here is the summary:
Generative AI (Artificial Intelligence, Large Language Models) offers the unprecedented ability tomanipulate and generate text and media in response to arbitrary instructions. These new capabilities offeropportunities and risks to researchers. This document will discuss responsible use, risk mitigation, andappropriate use of these tools. The technological landscape changes quickly and new tools are releasedalmost weekly – this guide offers general advice which should be applied thoughtfully.
Regardless of any tools or technologies used, now or in the future, everyone at Macquarie University isresponsible for ensuring their research meets the expectations of the Australian Code for the ResponsibleConduct of Research and the Macquarie University Code for the Responsible Conduct of Research(2018).
Generative AI must be used with caution, and its use is currently inappropriate in some researchprocesses because Generative AI services, including ChatGPT:• cannot meet the requirements for authorship• can create authoritative-sounding outputs that may be incorrect, incomplete, or biased• could inappropriately capture sensitive data (including, but not limited, to personal information).
Researchers must not use Generative AI:• to perform peer review activities• to generate substantive content of research outputs, including HDR theses• for writing the critical components of human ethics, animal ethics, or biosafety applications.
Researchers must exercise care in the use of Generative AI in other aspects of their research and should:i. only do so with the written agreement of their research collaborators (& HDR supervisors)ii. review and consider the terms of service/license of the platforms used and any models usediii. consider the current issues and understandings around copyright and intellectual propertyiv. mitigate risks around the insecure storage or unauthorised re-use of sensitive datav. exert oversight and control when using the technologyvi. carefully and critically review the output and results created by Generative AIvii. take responsibility for the integrity of the content altered or created using Generative AIviii. disclose the use of Generative AI to potential publishers and in disseminated research outputsix. read and follow the policies of publishers and funders regarding the use of Generative AI.