INTS1302 Navigating Digital Worlds, Macquarie University
2026-08-11
The National Statement:
Human research is research conducted with or about people, or their data or their biospecimens.
Kaufman, defending Harvard’s release of student Facebook data:
Would you require that someone sitting in a public square, observing individuals and taking notes on their behavior, would have to ask those individuals’ consent in advance? We have not accessed any information not otherwise available on Facebook. We have not interviewed anyone, nor asked them for any information, nor made information about them public…
(Kaufman 2008, comment on (Zimmer, 2008))
The National Statement:
Research is ethically acceptable only when its potential benefits justify any risks involved in the research.
Cegłowski:
In 2007 LiveJournal is sold to a Russian company, and a few years later, to everyone’s surprise, homophobia is elevated to state ideology.
Kramer and colleagues, for the mood experiment:
We show, via a massive (N = 689,003) experiment on Facebook, that emotional states can be transferred to others via emotional contagion, leading people to experience the same emotions without their awareness.
Boellstorff:
The principle of care arises in part from asymmetrical power relations and imbalance of benefit between investigator and investigated. The investigator generally gains far more than the informant, garnering benefits that translate to jobs, money, and professional recognition.
Haraway:
Relativism is a way of being nowhere while claiming to be everywhere equally.
Markham:
platform and service providers assert that their algorithms are impartial, functioning through rule-based calculations on objective data points without any “subjective” interference.
Haraway:
I am arguing for the view from a body, always a complex, contradictory, structuring, and structured body, versus the view from above, from nowhere, from simplicity. Only the god trick is forbidden.
The Nuremberg Code:
The voluntary consent of the human subject is absolutely essential.
(“The Nuremberg Code (1947),” 1996, p. 1)
The National Statement:
On rare occasions the practice of research has even involved the deliberate and appalling violation of human beings, notoriously, the Second World War experiments in detention and concentration camps.
Facebook, after the mood experiment:
Although this subject matter was important to research, we were unprepared for the reaction the paper received when it was published and have taken to heart the comments and criticism. It is clear now that there are things we should have done differently.
Hu:
The FTC action, however, has been criticized as failing to adequately address the privacy and other harms emanating from Facebook’s release of approximately 87 million Facebook users’ data, which was exploited without user authorization.
Ms Claybaugh, for Meta:
Correct. We use data that people have expressly made public
(Adopting Artificial Intelligence (AI), 2024, p. 2)
A Facebook user:
…With stuff like the recent Facebook scandal, it’s like you don’t realise how open your data is. I feel like a lot of companies probably do have my data now and I’ve just kind of got to the point where I’ve accepted, the basic data, I don’t care about sharing that with third parties anymore because I know most of them probably have it by this point.
Hu, on what the American regulator did:
The settlement, announced on 24 July 2019, included a record-setting $5 billion fine and an FTC Order to institute new privacy standards
The National Statement:
The National Health and Medical Research Council Act 1992 (Cth) (NHMRC Act) establishes NHMRC as a statutory body and sets out its functions, powers and obligations. Section 10(1) of the NHMRC Act requires the Chief Executive Officer (CEO) to issue human research guidelines… All the guidelines in this National Statement that are applicable to the conduct of research involving humans are issued by NHMRC in fulfilment of this statutory obligation.
The NTIA:
Therefore, the secondary use of only non-identifiable data in research, for example, would generally not be subject to the Common Rule’s requirements, even for research that is federally supported or conducted.
Ethics review, the National Statement:
Importantly, the opt-out approach is unlikely to constitute consent when applying Commonwealth privacy legislation to the handling of sensitive information, including health information. Therefore, where it is impracticable to obtain an individual’s explicit consent to the use of their information… researchers must comply with the Guidelines under Section 95 of the Privacy Act 1988.
Research integrity, the Australian Code:
This Code does not incorporate the laws, regulations and guidelines and other codes of practice that apply to the conduct of research. Those responsible for the conduct of research are expected to be aware of and comply with the applicable laws and codes.
Privacy law, the OAIC:
Just because data is publicly available or otherwise accessible does not mean it can legally be used to train or fine-tune generative AI models or systems
Platform terms, Kogan’s defence:
Kogan contended that he conformed to Facebook’s guidelines at the time
boyd and Crawford:
Just because content is publicly accessible does not mean that it was meant to be consumed by just anyone.
(Boyd & Crawford, 2012, p. 12)
Nissenbaum:
The notion that when individuals venture out in public, a street, a square, a park, a market, a football game, no norms are in operation, that “anything goes,” is pure fiction.
The National Statement:
Unless a waiver of the requirement for consent is obtained, any research access to or use of publicly available data or information must be in accordance with the consent obtained from the person to whom the data or information relates.
The AoIR guidelines:
User-generated content is generally published in informal spaces that users often perceive as private but may strictly speaking be publicly accessible. In any case, researchers are rarely the intended audience of user-generated content.
The National Statement:
The guiding principle for researchers is that, although data or information may be publicly available, this does not automatically mean that the individuals with whom this data or information is associated have necessarily granted permission for its use in research.
al-Zaman and colleagues:
we believe that social media content, being publicly available and created voluntarily by users before the study, is free from ethical restraints
Ms Claybaugh, for Meta:
That means when you go on Facebook or you go on Instagram and you make a post, you select the audience for that post—that statement, that photo, whatever it is you’re posting online. If you choose to make that post, the text or the image, public, that is publicly sharing that information.
A Facebook user:
We all are very exposed aren’t we, in so many ways. Everything we do is exposed and we sell it to ourselves because most of the time they say, ‘Oh if you like us and comment here and give us your Facebook, you will get a free cappuccino from Starbucks’… so everyone’s going to give you all the data you want… Most people don’t care about their privacy… and they don’t even read the terms and conditions, they just click ‘accept.’
The OAIC, 21 October 2024:
a failure of a website to implement measures to prevent data scraping should not be taken as implied consent
Ms Claybaugh, for Meta, 11 September 2024:
We do provide an opt out so that people can say, ‘I no longer want my public posts and images to be used to train the models,’ so we do provide that opt out.
(Adopting Artificial Intelligence (AI), 2024, p. 8)
Senator Shoebridge, to Meta:
she would never have contemplated that Meta was going to scrape those photos … and yet you chose to just sweep all that information into your AI model without asking her consent. Can’t you see the ethical problem there?
Townsend and Wallace:
Of particular concern is the republishing of quotes that have been taken from social media platforms and republished verbatim, as these can lead us, via search engines, straight back to their original location, often then exposing the identity and profile of the social media user they originate from.
(Townsend & Wallace, 2016, p. 7)
Taylor and colleagues:
Google immediately identified the social media influencer in all de-identified trials, despite the anonymization techniques we had used. This indicated that traditional anonymization strategies for visual print data were, in fact, not effective in an online context.
The Berkeley Protocol, page 90:
Anonymization: the process of making it impossible to identify a specific individual.
(OHCHR & Human Rights Center, 2022, p. 90)
The Berkeley Protocol, page 25:
Investigators should also be aware of the mosaic effect, whereby public data, even when anonymized, may become vulnerable to reidentification if enough data sets containing similar or complementary information are released or combined.
Waldek and colleagues:
Additional safeguarding around the process of engagement with the data were also put in place, including time limits for viewing the data, mental health first aid assessment processes and trauma informed practice techniques.
The AoIR guidelines:
when is it allowable (if ever) to use data that would otherwise be prohibited ethically and/or legally because of privacy protections, etc., but has been made public because of an accidental breach and/or intentional hack
The National Statement:
This includes avoiding the use or disclosure of information that was obtained unethically or illegally.
Post:
to use [hacked] data without the consent of those who were violated is to violate the violated anew
(Post 1991, quoted in (Waldek et al., 2025, p. 13))
Poor and Davidson:
we want this data but we don’t need it
(Poor and Davidson 2016, quoted in (Waldek et al., 2025, p. 11))
An open-source journalist:
The importance of the investigative goal is often greater than the legality of the means.
Cegłowski:
In particular, I’d like to draw a parallel between what we’re doing and nuclear energy, another technology whose beneficial uses we could never quite untangle from the harmful ones.
Cegłowski:
A singular problem of nuclear power is that it generated deadly waste whose lifespan was far longer than the institutions we could build to guard it.
The AoIR guidelines:
Once an AI system has left the hands of the original researchers, they may not have any control over how their models are used by others. The same is true for the generated research data: once it has been freely published, it will be difficult to contain its further uses.
The European Commission:
Rather, the ORD pilot follows the principle “as open as possible, as closed as necessary” and focuses on encouraging sound data management as an essential part of research best practice.
The NHMRC’s data guide:
for areas such as gene therapy, research data must be retained permanently (e.g. data in the form of patient records).
The National AI Plan:
AI models are only as good as the data they are trained on… large, unstructured datasets could be made accessible for AI system training.
Van der Woude and colleagues:
Open source investigators’ narratives reveal that their work relies on informal peer control and their own personal adherence to the “do not harm” and “public interest” principles rather than a professionally and widely agreed code of ethics.
(Van Der Woude et al., 2025, p. 14)
The Berkeley Protocol:
Open source investigators must be accountable for their actions, which can often be ensured through clear documentation, record-keeping and oversight.
(OHCHR & Human Rights Center, 2022, p. 24)
The National Statement:
The researcher is responsible and accountable to their institution, any sponsors or funders of the research, participants and, in some research, to regulators or other entities who have a formal role in the oversight of the research.
An open-source journalist:
We don’t have any guidelines, and I think that’s crooked: I wish we did have rules. We do talk about privacy considerations with colleagues continuously, but we don’t have standardized rules, and I think we should have them.
The AoIR guidelines:
the issues raised by Internet research are ethical problems precisely because they evoke more than one ethically defensible response to a specific dilemma or problem. Ambiguity, uncertainty, and disagreement are inevitable.
The National AI Plan:
Australia has strong protections in place to address many risks, but the technology is fast-moving and regulation must keep pace. That’s why the government continues to assess the suitability of existing laws in the context of AI.