Latest
Daily Tech Times Subscribe

Data and IP ownership in AI startups: what UK founders should get right early

There is currently no general exception in UK law letting an AI company freely use copyright-protected material to train commercial models, and the government has deliberately deferred a final decision on reform, so UK AI founders are working with more legal uncertainty here than many assume.

Yellow and green cables are neatly connected
Photo · Photo by Albert Stoynov on Unsplash

This is a general, educational overview of the current legal landscape for data and IP in UK AI startups, not legal advice, and the underlying law is genuinely in flux, so anything specific to a particular company’s data sources, model architecture or customer contracts needs proper advice from an IP or technology solicitor. The short version is that UK copyright and data protection law were not written with AI training in mind, the government has been actively consulting on reform since late 2024, and as of the most recent government report on the subject, it has deliberately deferred a final decision rather than legislating a clear answer.

Can a UK AI company legally use copyrighted material to train its models?

Not freely, and not without either a licence or a specific applicable exception. Unlike the US, where a broader “fair use” doctrine gives AI developers more room to argue that training use is permitted, UK law relies on a narrower set of specific exceptions, and there is currently no general text and data mining exception that covers commercial AI training on copyright-protected material without permission. The government consulted between Decemberundefinedand Februaryundefinedon introducing a broader exception with an opt-out mechanism for rights holders, but its own subsequent report on the consultation confirmed it has rejected that as its preferred option, at least for now, citing strong opposition from creative industries and concerns that requiring individual creators to actively “opt out” would be unfairly burdensome, particularly for smaller rights holders. The practical result is that a UK AI startup training on third-party content generally needs to either licence that content directly, rely on a genuinely applicable existing exception (which is a narrower fit than many assume), or use content it can otherwise establish it is legally entitled to use, such as its own first-party data or content properly licensed for this purpose.

Is this settled law, or is it likely to change?

It is not settled, and founders should treat it as an area to actively monitor rather than a fixed rulebook. The government’s own report states it will not introduce copyright reform “until we are confident that they will meet our objectives,” and describes further evidence-gathering and monitoring of how other jurisdictions handle the same question as the near-term plan rather than firm UK legislation. One area where the government has been clearer is proposed reform to section 9(3) of the Copyright, Designs and Patents Act 1988, which currently gives a limited form of copyright protection to “computer-generated works” with no human author; the government has indicated it intends to move away from protecting wholly AI-generated output this way, while preserving protection for AI-assisted work that still involves genuine human creative input, on the basis that copyright should incentivise human creativity specifically. Because this is genuinely a live policy area, a startup relying on the current position for training data sourcing, or on the current position for owning AI-assisted output, should build in a periodic legal review rather than assuming today’s position is permanent.

Database right is a distinct UK intellectual property right, separate from copyright, that protects a database where there has been substantial investment in obtaining, verifying or presenting its contents, regardless of whether the arrangement of that content shows the kind of originality copyright requires. This matters a great deal for AI and data-heavy startups because a company’s proprietary training data set, a curated collection of labelled examples, scraped and cleaned records, or aggregated customer data, may be protectable under database right even where individual pieces of the underlying data are not themselves copyright-protected. Ownership of database right belongs to the “maker”, generally the person or organisation that took the initiative and bore the financial risk of creating the database, which is a genuinely different test from who owns copyright in a work, and it means a company using an external contractor to build or enrich a dataset should not assume database right, any more than copyright, automatically transfers to the company without a clear written agreement.

What does GOV.UK’s general guidance say about IP created for a business?

GOV.UK’s general “IP Basics” guidance makes a point that catches out a lot of early-stage companies: if you commission a third party, a freelance developer, a data labelling contractor, an agency, to create copyright-protected work for your business, ownership does not automatically transfer to you unless this is agreed in writing before the work is created. In practice this means an AI startup that has used contractors or agencies to build training pipelines, label data, or develop model code needs a clear IP assignment in every relevant contract, ideally in place before work starts rather than negotiated retrospectively, since a company that skips this step can find a contractor retains rights that materially limit how freely the company can use or commercialise what was built.

What does data protection law require when personal data is involved?

Where an AI system is trained on or processes personal data, the UK GDPR framework applies in full, and the Information Commissioner’s Office’s guidance on AI and data protection confirms that existing obligations, having a lawful basis for processing, data minimisation, purpose limitation, keeping data accurate, and accountability, apply across the AI lifecycle from initial data collection through model training to deployment and ongoing monitoring, rather than there being some separate, lighter-touch regime for AI specifically. For a startup, this means the lawful basis question, on what legal ground is personal data being collected and used for training, has to be worked out and documented at the point data is first gathered, not retrofitted once a model already exists, since the ICO’s guidance treats each stage of the AI lifecycle as needing its own genuine assessment rather than a single blanket justification covering everything downstream.

What should a founder actually do early on?

It is worth mapping out, honestly, where every significant training data set actually came from, whether it is owned outright, properly licensed, obtained under a genuinely applicable exception, or of uncertain provenance, since this map becomes exactly what an investor’s technical and legal due diligence will probe before a funding round, and gaps discovered late are far more costly to fix than gaps identified and addressed early. It is equally worth putting in place, from the very first contractor or freelancer engagement, a standard IP assignment clause covering anything built for the company, code, datasets, models, and a data protection framework that documents the lawful basis for any personal data used, rather than treating either as something to formalise only once the company is further along and facing investor questions.

The practical takeaway

UK data and IP law for AI is genuinely unsettled in places, particularly around training data and copyright, which is a real source of risk that a founder should acknowledge honestly rather than assume is resolved simply because AI products are now widespread. Because the specific facts of a company’s data sources, contracts and model architecture determine its actual exposure, and because the underlying law is actively being reviewed by government, it is worth getting a solicitor experienced in technology and IP law to assess a company’s specific position, and to flag it for periodic review, rather than relying on a general educational overview such as this one.

Sources