You have four years of unpublished research, a folder of interview transcripts collected under an ethics approval, possibly an industrial partner’s confidentiality agreement, and a text box that will happily accept all of it. The question is not paranoid and it is not rhetorical: doctoral researchers hold the most sensitive material of any student group, and almost nobody has told them where the line is. Tesify is a workspace built to hold a whole thesis, not a chatbot you paste into, and it is free to start — but the point of this page is the reasoning, because you will use more than one tool over four years and you need a rule, not a recommendation.
This page is about your research material: participant data, unpublished results, and work under a confidentiality agreement. It is not about language editing or authorship, which are a separate question covered in our guide to AI academic English editing, and it is not about proofreading services.
Start with what the law actually says, because the common claim is wrong
The claim circulating in most PGR common rooms is that data protection law forbids putting participant data into a third-party tool. That is not what it says, and getting the rule right matters because the true version is both more permissive and more demanding.
The Information Commissioner’s Office sets out the research provisions plainly. Article 89 of the UK GDPR makes use of those provisions dependent on “appropriate safeguards” taking “the form of technical and organisational measures”, and Article 89 “specifically mentions measures to ensure respect for the principle of data minimisation. This may involve, where possible, anonymising or pseudonymising data.”
Then comes the sentence that reframes the whole question. On anonymisation, the ICO states: “Anonymous information is not personal data. Data protection law does not apply.”
So the requirement is not a prohibition on tools. It is a sequence. Data minimisation under Article 5(1)(c) means personal data should be “adequate, relevant and limited to what is necessary in relation to the purposes for which they are processed”, and the ICO’s instruction is to work back from that: “You should first consider whether it is possible to conduct your research without using personal data. If you could carry out your research using anonymised data, then you do not need to process personal data.”
Where you cannot anonymise, pseudonymisation is the fallback — and the ICO is careful that it is not the same thing: “Pseudonymous data is still personal data and data protection law applies.” Either way, timing matters: “you should ensure that you do anonymisation or pseudonymisation at the earliest possible opportunity, ideally prior to using the data for research purposes.”
Two hard limits sit on top. Section 19 of the Data Protection Act 2018 states that research-related processing does not satisfy Article 89 if it “is likely to cause people substantial damage or substantial distress” or “is carried out for the purposes of measures or decisions about particular people”, except for approved medical research.
One currency note, because this is exactly the kind of claim that goes stale: the ICO’s own research-provisions guidance currently carries a banner stating that, due to changes made by the Data (Use and Access) Act, the guidance is under review and may be subject to change. Read the current version rather than a summary — including this one — before relying on it for a decision that matters.

The four categories, and only one of them is genuinely hard
1. Your own prose. Your draft chapters, your argument, your literature synthesis. No personal data, no third-party confidentiality, no participants. The only questions here are the tool’s retention terms and your institution’s authorship rules, and both are manageable.
2. Anonymised or aggregated results. Summary statistics, model outputs, coded themes with no identifiers. On the ICO’s own statement, anonymous information is not personal data and data protection law does not apply — so the analysis here is about commercial confidentiality and prior publication, not about data protection.
3. Identifiable participant data. Raw interview transcripts with names, places of work and clinical detail; recordings; anything from a small population where deductive disclosure is possible. This is the hard category, and the rule is the sequence above: anonymise or pseudonymise first, at the earliest opportunity, and only then decide what tool touches it.
4. Third-party confidential material. Data belonging to an industrial partner, an NHS trust or a sponsoring organisation, governed by an agreement you probably signed at the start and have not read since. This category is not governed by data protection reasoning at all — it is governed by a contract, and the contract may prohibit disclosure to any third party regardless of whether personal data is involved.
What your ethics approval already committed you to
This is the constraint doctoral researchers forget, and it binds independently of the law. Your ethics application described how data would be stored, who would have access to it and how long it would be kept. Your participant information sheet told people that, and they consented on that basis.
If your consent form says the recordings will be held on university-managed storage and accessed only by the research team, then routing them through a transcription service or a chatbot is a departure from what your participants agreed to — whatever the tool’s privacy policy says. That is a research-integrity question before it is a legal one, and the remedy is not to reason about it privately but to go back to your ethics committee for an amendment if the workflow has changed. Committees approve amendments routinely; they take a much dimmer view of discovering the change afterwards.
This bites hardest on audio. Recorded interviews are the most identifiable material most doctoral researchers hold — a voice is a biometric identifier and a transcript may name colleagues, employers and patients. Our comparison of transcription tools for PhD research interviews goes through the accuracy, ethics and cost trade-offs specifically, and the data-handling question there is the same one as here, one step earlier in the pipeline.
The other two risks nobody frames as data protection
Your NDA. If your project is industry-sponsored or embedded in an organisation, read the agreement before you read any privacy policy. Confidentiality clauses commonly prohibit disclosure to any third party without written consent, and a cloud tool is a third party. It does not matter that the data is anonymised or that the vendor promises not to train on it. Ask your supervisor or your university’s research contracts team whether the agreement permits third-party processing at all; that is a two-email question with a written answer.
Prior publication and embargo. Journals care about whether material has been publicly disclosed before submission. Pasting a passage into a tool with confidential-input terms is not publication; posting it into a public forum, a shared community workspace or anywhere indexable can be. If you intend to publish a paper from your thesis or to submit a thesis by publication, keep unpublished material out of anything with a public surface, and keep the embargo question in view from the start rather than at deposit.

Five questions to ask of any tool, before you upload anything
- Is my input used to train the provider’s models, and can I turn that off? Consumer and professional tiers of the same product frequently differ on exactly this. The answer lives in the terms, not the marketing.
- How long is my content retained, and can I delete it? “We do not train on your data” and “we delete your data” are different promises, and vendors often make only the first.
- Where is it processed, and who else can see it? Sub-processors, staff review for abuse monitoring, and jurisdiction all matter if your ethics approval or your NDA specified storage arrangements.
- Is there a written agreement my university would accept? Where personal data is involved, an institution acting as controller needs a proper processing arrangement, not a consumer click-through. Your research office knows which tools already have one.
- Can I export everything and leave? A four-year project should never be locked inside a service, and export is also your insurance against a vendor changing its terms mid-candidature.
Two practical shortcuts. Your university almost certainly has an approved-tools list and a research data management policy; checking them takes ten minutes and answers question four for free. And if you can answer these five questions out loud about every tool in your workflow, you can also answer them in a viva, which is the standard worth holding yourself to — the same test our guide to the notes-to-chapter writing-up workflow applies to authorship.
A workflow that is defensible end to end
- Anonymise or pseudonymise at collection, not at analysis. The ICO’s “earliest possible opportunity” is a genuine instruction, and it also makes every later decision easier.
- Keep the identifier key separate, on university-managed storage, and never let it near a third-party service.
- Do analysis on the de-identified set. If your method needs the identifiers, that is a design question for your ethics application, not a tooling question.
- Choose one workspace for the writing and stay in it. The risk in most doctoral workflows is not one careless upload; it is fragmentation — the same chapter scattered across four services with four different retention policies, none of which you have read.
- Record what you did. A short note of which tools you used, for what, and under what settings. It costs nothing, and it turns an awkward viva question into a two-sentence answer.
That fourth point is the one Tesify is built around. Rather than pasting fragments of your thesis into whatever tool is open, the chapters, structure and bibliography live in one workspace, and your work stays yours — not published, not shared into any public database, and exportable whenever you want it. Everything remains 100% written by you: the research, the reasoning and the words. Start free and keep the whole thesis in one place — which matters most across the long candidatures our page on how long a UK PhD actually takes describes, where a four-year trail of scattered drafts is the real exposure.
For choosing the discovery and appraisal tools that sit earlier in the pipeline, our comparison of the best AI research assistants for PhD students covers what each does with your queries and your library.
Frequently asked questions
Is it against UK data protection law to put my research data into an AI tool?
Not as a blanket rule. The ICO’s position is that anonymous information is not personal data and data protection law does not apply to it, so the first question is whether you can anonymise. Where you cannot, pseudonymised data is still personal data and the safeguards under Article 89 apply.
What does “appropriate safeguards” actually mean?
Technical and organisational measures, with Article 89 specifically naming data minimisation — which the ICO says “may involve, where possible, anonymising or pseudonymising data”. Section 19 of the DPA 2018 adds that research-related processing does not satisfy Article 89 where it is likely to cause substantial damage or substantial distress, or where it is for measures or decisions about particular people, except in approved medical research.
Is pseudonymised data safe to upload?
It is still personal data, so the same obligations apply as to any other personal data — it is a risk reduction, not an exemption. Anonymisation is the step that takes the material outside data protection law entirely.
Does my ethics approval restrict which tools I can use?
Very likely, because your application described how data would be stored and who would access it, and your participants consented on that basis. If your workflow has changed, submit an amendment rather than proceeding and explaining later.
Can I put interview recordings into a transcription service?
Only if your ethics approval and consent documentation cover it, and only after checking the service’s retention terms. Audio is the most identifiable material most doctoral researchers hold, which is why it deserves the most careful handling in the whole pipeline.
My project has an industrial NDA. Does anonymising solve it?
No. Confidentiality agreements typically restrict disclosure to third parties regardless of whether the material contains personal data, so the analysis is contractual rather than data-protection-based. Ask your research contracts team for a written answer.
Will using an AI tool count as prior publication?
Not where the tool treats your input as confidential. The prior-publication risk comes from public disclosure — forums, shared public workspaces, anything indexable — so keep unpublished material off public surfaces if you intend to publish from the thesis.
How do I know whether a tool trains on my content?
Read the terms rather than the homepage, and check whether the setting differs between tiers of the same product. If you cannot find a clear answer, treat that absence as the answer for anything sensitive.
Do I have to declare tool use to my examiners?
Follow your institution’s disclosure requirements, which increasingly extend to AI used inside editing tools. A short, accurate statement of what you used and for what is cheap, and it removes the question from the room.
Is Tesify safe for unpublished doctoral research?
Your work in Tesify is yours: it is not published or shared into any public database, and you can export it at any time. As with any tool, keep your own exported backups, and apply the anonymisation sequence above to participant data before it goes anywhere.
Who at my university can give me a definitive answer?
Your data protection officer or research office for the data protection question, your ethics committee for the consent question, and research contracts for anything under an NDA. All three answers are worth having in writing before you build a workflow around a tool.
