Categories
Uncategorised

Media Provenance for Authenticity and Creative Content Exchange 

Speaker: Professor John Collomosse, University of Surrey

Abstract: Countering modern disinformation requires moving beyond AI detection toward embedded trust signals that travel with content. In this talk, I present media provenance as an approach for encoding verifiable information about origin and modification directly into media assets.  I will describe our work on embedding durable provenance signals using digital watermarking, content fingerprinting, and cryptographically signed metadata (C2PA) enabling provenance information to persist across common transformations and distribution channels. Beyond its role in content authenticity, I will show how provenance can also underpin new models for consent and value creation in the age of generative AI, including the use of causal data attribution and decentralised registries to ensure recognition and reward for creators and collection owners whose work contributes to model training. 

Bio: Professor John Collomosse is the founder and director of DECaDE, the UKRI Research Centre for the Decentralised Digital Economy led by the University of Surrey. He is a Senior Principal Scientist at Adobe Research where he directs Adobe’s research programme for the Content Authenticity Initiative (CAI), and two cross-industry task forces within the Coalition for Content Provenance and Authenticity (C2PA) standards body.  John’s research intersects Artificial Intelligence (AI) and Distributed Ledger Technology (DLT), with focus on media provenance to fight misinformation and online harms, and on improving data integrity and attribution for responsible AI.  He holds a PhD in Computer Science from the University of Bath, and is a fellow of the IET, BCS and Royal Society of Arts.  He is a member of the UKRI SETB advisory group, DCMS Creative Industries Council, and from 2008-2014 was a member of the strategic advisory teams (SATs) for the EPSRC ICT and UKRI Digital Economy programmes. 

Categories
Uncategorised

Balancing Access and Risk: Sensitivity-Aware Search in Sensitive Archives 

Speakers: Professor Iadh Ounis & Dr Graham McDonald, University of Glasgow

Abstract: Organisations such as governments, archives, and cultural heritage institutions are typically required to provide public access to large document collections in the interest of transparency, accountability, and research. However, such collections can often contain sensitive information such as personal data, confidential communications, or security-relevant information that cannot be safely released to the public. Traditional manual sensitivity review processes for removing sensitive documents prior to public release are expensive, time consuming, and can be difficult to scale. Therefore, there is a need for technical solutions that support safe and responsible search over such collections while also protecting any sensitive information, so that such information is not leaked into the search results. In this talk, we will present our recent work on developing AI‑powered technologies to help organisations safely release large collections of documents to the public without the need for a costly manual sensitivity review. In particular, our talk will primarily focus on our approaches to sensitivity‑aware search, which apply sensitivity‑filtering interventions at each stage of a modern search pipeline. Rather than treating sensitivity detection as a single classification problem, our work explores how safeguards can be embedded within the ranking models that are deployed throughout the search process. These safeguards reduce the risk of exposing sensitive information while still allowing the public to search and use the documents effectively. We will discuss how these approaches can help to balance the competing goals of maximising access to information while minimising disclosure risks. We will also present a new freely available document collection that we have created to support researchers and practitioners. This resource is designed to help develop and evaluate AI systems for sensitivity classification and sensitivity‑aware search, with the goal of encouraging research in this important area and ultimately improving the safe and responsible public access to sensitive archives.   

Bios

Graham McDonald 

Graham McDonald is a Senior Lecturer in Information Retrieval at the University of Glasgow. His research focuses on technology‑assisted sensitivity-aware retrieval and classification, safe information extraction, and responsible information retrieval for complex document collections. After completing his PhD in 2019 on Technology‑Assisted Sensitivity Review, which developed technologies for open government in the era of digital records, he led a number academic–industry knowledge exchange projects with SVGC Ltd. and FCDO Services to support the FCDO’s efficient execution of Digital Sensitivity Review through tailored state-of-the-art digital solutions. He was Principal Investigator of the EPSRC New Horizons project Safe Information Extraction from Patient Histories (SIEPH) and is currently a co-Investigator on the RAI UK project Participatory Harms Auditing Workbenches and Methodologies (PHAWM). Graham has held several significant leadership roles in the research community. He served as General Co‑Chair of ECIR 2024, Finance Chair for IEEE ICDCS 2025, co‑organised the NIST TREC Fair Ranking Track, delivered the Search Among Sensitive Content tutorial at ECIR 2021 and is co-organiser of the AI & Open Government Workshop at ICAIL 2026. He is an active member of the Information Retrieval community, serving on SPC and PC committees for leading conferences and journals, and has published more than 30 papers in these venues. He is also Vice Chair of the BCS Information Retrieval Specialist Group.   

Iadh Ounis 

Iadh Ounis is Professor of Information Retrieval in the School of Computing Science at the University of Glasgow. His research focuses on developing artificial intelligence technologies that support effective access to information, including advances in search engines, recommender systems, and conversational assistants. He is the principal investigator of the Terrier open‑source Information Retrieval platform and a contributor to its Python‑based AI extension, PyTerrier, widely used in academia and industry to support experimentation in AI‑driven search. A recipient of the UKeiG Tony Kent Strix Award, he has published over 300 papers in leading computer science venues, led several major international research initiatives, co‑chaired major international conferences including ECIR 2024 and IEEE ICDCS 2025, and led numerous UKRI, EU, and industry‑funded projects. His recent work includes contributions to AI‑supported sensitivity review for archives, collaborating with public sector and industry partners to help organisations manage and release information safely and responsibly. He has served in senior leadership roles within the Scottish Informatics and Computer Science Alliance (SICSA) and The Data Lab, the Scottish Funding Council Innovation Centre for data and AI, and is currently a member of The Guild’s EU Heads of Artificial Intelligence and Digital Research Group. 

Categories
Uncategorised

A Use Case for Agentic AI Enhancing Access to Public Records: Challenges & Opportunities 

Speaker: Professor Jason R. Baron, University of Maryland

Abstract: Given ever-increasing, enormous volumes of public records in electronic form, many government departments and agencies struggle to meet their disclosure obligations under freedom of information (FOI) laws. The challenges public authorities face are magnified due to existing backlogs and ongoing fiscal and resource constraints, all of which contribute to substantial delays in processing new requests. In the United States, the Freedom of Information Act (FOIA) sets a statutory time frame of 20 working days for federal agencies to respond to requests, yet government lawyers increasingly have been called on to file declarations under oath from agency officials that it may take decades to complete a large volume request response. The introduction of artificial intelligence (AI) technologies into access request workflows holds out the potential to dramatically improve the FOI “user experience,” which we will refer to here as the “requester’s experience” or “RX.” This lecture will first survey how federal agencies in the United States are employing AI tools to improve processing of requests filed under the FOIA, including using machine learning and chatbot capabilities. The lecture will proceed to describe how agentic AI might transform a future FOIA process for requesters into a seamless, automated, interactive, end-to-end workflow, subject to human-in-the-loop quality controls and supervision at key stages. Notwithstanding legitimate concerns regarding the future use of “black box” algorithmic processes, the view will be expressed that on balance using AI methods will lead to a more transparent and accountable government. 

Bios: Jason R. Baron is a professor of the practice in the College of Information at the University of Maryland. During his 33 years in government, he served as the first appointed director of litigation at the US National Archives and Records Administration, and before that as a trial attorney and senior counsel at the Justice Department. In those capacities, Mr. Baron acted as lead counsel on landmark lawsuits involving the preservation of White House email and played a leading role in improving federal electronic recordkeeping policies. He is a recipient of the international Emmett Leahy Award, given for achievements in records and information management. Mr. Baron received his B.A. magna cum laude with honors from Wesleyan University, and his J.D. from Boston University School of Law.     

Categories
Uncategorised

Smarter Government Records: Using Knowledge Management and AI to Unlock Hidden Value  

Speaker: Nicholas Smith, NHS Resolution

Abstract: Government bodies hold vast collections of records that document decisions, policies and public services, yet much of their hidden value remains under exploited due to fragmented knowledge practices and legacy information systems. This presentation explores how modern knowledge management approaches, combined with responsible, outcome focused applications of artificial intelligence, can transform these records into a strategic asset for better governance, improved operational efficiency, enhanced public accountability and greater confidence in decision making. 

Drawing on practical experience from NHS Resolution’s evolving Knowledge and Information Management function, the session examines the shift from traditional records centric models to approaches that emphasise organisational knowledge, context and reuse. It will look at how establishing clear frameworks for capturing key knowledge, designing consistent metadata, and embedding knowledge capture activities throughout the records lifecycle provides the essential foundations for both AI readiness and exploitation. With these foundations in place, the presentation will illustrate how AI tools with the support of knowledge management can support discovery, synthesis and insight generation across large, complex and historically underutilised records, while maintaining the ethical safeguards necessary for use in the public sector and ensuring compliance. 

Real world examples will show how aligning knowledge management principles with AI techniques can streamline workflows, improve established memory and unlock innovative use cases, from surfacing patterns in historic casework to enhancing transparency through richer, more accessible public records. The session concludes by outlining practical steps government organisations can take today to build knowledge management and AI maturity, ensuring that public records are not only preserved but actively leveraged to deliver measurable public value and efficiency. 

Bio: I am experienced leader in the public sector with a strong background in governance, knowledge, records, information and project management. As the Head of Knowledge and Information Management at NHS Resolution, I provide strategic leadership and oversees the records, knowledge, and information management functions. Also, previously having worked for the MOD, HM Treasury, Cabinet Office and the Department for International Trade. During my career specialising in records and retention, knowledge and information management, governance, setting up and establishing records, knowledge and information functions and identifying how AI can be used with to enhance records, knowledge and information management. 

Categories
Uncategorised

Public AI for Public Archives: Collaborative R&D in UK Screen Archives through the ISSA project 

Speaker: Dr Daniel Chávez Heras, King’s College London

Abstract: When archives purchase AI services, they receive outputs—automated transcription, entity extraction, metadata generation—but the knowledge of how those outputs were produced remains with the vendor. Each token bought is capacity not built. This transactional model offers efficiency but forecloses institutional learning: archives become consumers of AI rather than participants in its development, and the sector’s collective understanding of what these technologies can and cannot do remains shallow. 

The Intelligent Systems for Screen Archives (ISSA) project, funded by the BFI through the National Lottery, explores an alternative path. Five UK regional and national moving image archives are collaborating with King’s College London to develop open-source, modular prototypes for metadata enrichment, collection visualisation, and retrieval. The investment is slower and harder than direct procurement, but the returns are different in kind: shared infrastructure, transferable methods, and crucially, collective knowledge about AI’s capabilities and limitations in archival contexts. 

ISSA’s approach alternates between engagement and development: fifteen interviews with archive professionals were distilled into four use cases and technical requirements, followed by iterative prototyping toward minimum viable products, a demonstrator event with all partners, and upcoming situated workshops that will apply these tools to specific archival challenges across the UK. This structure is designed not only to produce functional tools, but to generate and circulate practical knowledge across institutions with different scales, capacities, and priorities. 

This presentation reflects on early lessons from ISSA’s first phase and shares the project’s direction of travel. While ISSA focuses on moving image collections, the underlying questions—about ownership, capacity-building, and the distribution of knowledge in public AI development—are shared across the GLAM sectors and resonate directly with the challenges facing government archives as they navigate AI adoption. The goal is not to offer a model to replicate, but to contribute one example to a broader conversation about what public investment in AI research and development might look like. 

Bio: Daniel Chávez Heras is Lecturer in Digital Culture and Creative Computing in the Department of Digital Humanities at King’s College London, and Principal Investigator of the ISSA project. His research combines critical frameworks from film, television, and media studies with technical practice in creative and scientific computing, including applied machine learning. He works with cultural institutions including the British Film Institute, the British Council, and the BBC, and is the author of Cinema and Machine Vision: Artificial Intelligence, Aesthetics and Spectatorship (Edinburgh University Press, 2024). 

Categories
Uncategorised

The Three Steps to Copilot 

Speaker: Robert Bath, Intelogy

Abstract: Rob will provide an overview of the three important steps that every government body should prioritise before they roll our Copilot. Specifically, he will discuss: 

  • An overview of tools for of securing your content. This will outline approaches for excluding sites from Copilot, the use of sensitivity labels and data loss prevention policies and the new unified Data Security Posture Management capability. 
  • The importance of providing staff with clear guidance around Copilot. Including reinforcing ethical usage of AI and the necessity of never using AI to cheat knowledge 
  • Why records management is essential for improving the quality of Copilot’s results. 

Rob will also present some examples where Copilot has generated incorrect results, to emphasise the importance for government departments and agencies of maintaining high-quality information in Microsoft 365. 

Bio: Having overseen some of the most significant records management solutions in Microsoft 365, Rob has become a well-known Information Governance expert who regularly provides his thoughts to the wider community. Currently serving as the IRMS’ Digital Director, Rob has recently launched Orinoco 365 a product that makes it much easier to consistently apply Microsoft’s retention capabilities to your content in Microsoft 365. 

Categories
Uncategorised

Big Data and AI 

Speaker: Dr Haibin Cai, Loughborough University 

Abstract: This talk provides a concise review of recent advances in big data and artificial intelligence, with a particular focus on the rapid development of large language models (LLMs) and their transformative impact across a wide range of domains. Key technical innovations will be highlighted, including Retrieval-Augmented Generation (RAG), which enables the integration of external knowledge into generative models, and ensemble strategies that leverage multiple smaller models to improve performance, robustness, and trustworthiness. 

Bio: Dr Haibin Cai received his PhD from the School of Computing at the University of Portsmouth in 2018. He is currently a Senior Lecturer in the Department of Computer Science at Loughborough University, a leading public research institution renowned for excellence in science and engineering. Prior to his current position, Dr Cai held research roles at the University of Portsmouth, Loughborough University, and Imperial College London. He has developed extensive expertise in computer vision, artificial intelligence, and deep learning through his involvement in a range of research projects, including the EU FP7-funded project on robot-enhanced therapy for children with autism spectrum disorders, and the Innovate UK-funded projects. His current research interests include large language models (LLMs), human motion recognition, multisensory sensing, defect detection, and real-time deep learning-based applications. 

Categories
Uncategorised

Trustworthy by Design: Balancing AI-Driven Access with Provenance, Sensitivity, and Safety in Government Records 

Speaker: Dr Deblina Bhattacharjee, University of Bath 

Abstract: As artificial intelligence becomes embedded in how government records are processed and made accessible, a central question emerges: how do we ensure that the pursuit of openness does not come at the cost of trustworthiness? This talk examines the tension between AI-driven access and the principles of provenance, sensitivity, and safety that underpin archival practice. 

Drawing on recent developments across institutions including the UK National Archives, NARA in the United States, among others, the presentation surveys how AI is already being applied- from semantic search and automated metadata generation to sensitivity screening and document transcription. It then turns to the risks these applications introduce: hallucinated descriptions entering the archival record, re-identification of anonymised individuals, and the erosion of provenance when AI-generated content becomes indistinguishable from the original. 

The talk also draws on my own work at the intersection of AI and cultural heritage, where I develop AI methods for analysing imagery such as artworks and illustrated archives. This experience has demonstrated a recurring lesson: the value of any AI output depends entirely on how well the underlying data is prepared, how transparently the process is documented, and whether domain experts remain central to interpretation. These principles, developed in the context of visual collections, apply directly to the challenges facing government records. 

Rather than proposing that AI should or should not be used, the presentation argues that trustworthy use depends on design choices made before deployment: investing in data readiness, maintaining human oversight, labelling AI contributions explicitly, and building governance frameworks that keep pace with the technology. The talk is intended to open a broader discussion among participants about what responsible AI practice looks like across different institutional contexts. 

Bio: Dr. Deblina Bhattacharjee is an Assistant Professor in Computer Science at the University of Bath, where she leads research at the intersection of AI and cultural heritage. Her work applies AI techniques, including segmentation, depth analysis, and style transfer, to imagery such as artworks and illustrated archives. She has created widely used datasets for analysing visual art and has organised international workshops bringing together computer scientists and cultural heritage practitioners. Her current research extends to both imagery and documents, using AI to inform questions of provenance and safety. She actively collaborates across disciplines to support the accessibility of cultural collections.  

Categories
Uncategorised

Machine‑Interpretable Selection Outline (MISO) 

Speaker: Chris Royds, The National Archives  

Abstract: The number of born-digital government records is already vast, and growing exponentially. Traditional paper-based approaches to managing these “digital heaps” are not viable at scale; some sort of machine assistance is needed. Furthermore, appraisal and selection decisions require an understanding of the context behind these records – which, in many cases, has been lost over time. 

The MISO project explores how the archive can help government departments address these gaps in their organisational memory by providing machine-readable contextual data derived from TNA’s existing collections – the ‘memory of government’. We hypothesise that TNA’s digital collections contain rich insights into past organisational structures, business priorities, initiatives, and policy development. In a structured, AI-ready format, these insights can inform algorithms and AI tools to map a department’s legacy digital records, providing historical context to guide selection decisions. We call this resource the Machine-Interpretable Selection Outline (MISO). 

In addition, several of the techniques used in gathering and summarising information from TNA’s digital collections may be useful in understanding the contents of government departments’ digital heaps themselves. Implemented in the right way, and placed in the hands of records managers, these techniques could significantly increase the pace of working through existing digital heaps. 

Bio: Chris Royds is a Data Scientist at The National Archives. His work focuses on developing data-driven techniques, to assist government in dealing with large heaps of unstructured digital records. He has previously worked at HMRC, The Met Office, Birkbeck University, and KPMG. He even had a small speaking role in the 2016 horror film The Conjuring 2. 

Categories
Uncategorised

De-redacting the Archive: LLM-based Redacted Entity Estimation

Speaker: Professor Reuben Binns, University of Oxford 

Abstract: Redaction of sensitive details from publicly released archives has been a common means of protecting privacy and secrecy for decades. Speculation about redacted identities has long been a feature of such archives, from the Cold War intelligence files to the Epstein e-mail archive today. Digital tools, including AI, have long been used to assist in the process of redaction, identifying references to personal data, entities, and sensitive details, and flagging them for human review. However, with the increasing capabilities and widespread adoption of Generative AI tools like ChatGPT, they might also be used to undermine the efforts to redact information from archives. 

This talk will assess some of the possible threats of AI-driven de-redaction. Traditional, manual methods of guessing redacted identities depend on strong historical knowledge, analysis of context, and even typographic details like the length of the redaction bar. Now, budding de-redactors are likely to turn to AI to aid in their guesswork. This talk will present a preliminary assessment of the capabilities of current AI models for this task, through a series of case studies. It will also consider the risks that might arise when AI is used for de-redaction. Risks could arise both when AI is effective at de-redaction, undermining balance between public interest and privacy that redaction aims to uphold; and also when they are ineffective, leading users down blind alleys and inducing potentially misleading, harmful, or even defamatory inferences. 

Bio: Reuben Binns is an Associate Professor of Human Centred Computing, working between computer science, law, and philosophy, focusing on data protection, machine learning, and the regulation of and by technology. Between 2018-2020, he was a Postdoctoral Research Fellow in AI at the Information Commissioner’s Office, addressing AI / ML and data protection. He joined the Department of Computer Science at the University of Oxford as a postdoctoral researcher in 2015. He received his Ph.D. in Web Science from The University of Southampton in 2015.