Categories
Uncategorised

Beyond the Hype: Building Solid Data Foundations for Knowledge Transformation

Speaker: David Canning MBE & Dr Kelcey Swain, Cabinet Office   

Abstract: AI adoption is a top priority for government and a subject filling column inches and TV schedules, but as information and records professionals we need to understand the specifics that lie behind the hype. David will discuss some of the projects that Cabinet Office has pursued and the challenges we have faced. These include document and email records review, implementation of Google Gemini, and experiments with solid data foundations before using AI. This includes working on an AI report writer for data subject access and how to construct policies in machine-readable formats using solid data structures. David will discuss how, without solid data foundations, AI can often be a new problem rather than a solution. However, if implemented well, AI can transform access to knowledge and information into structured, machine-readable data, improving access through the conversion of knowledge into human-readable and human-owned data. 

Bio:

David Canning MBE 

David Canning is Head of Digital Knowledge & Information Management in Cabinet Office.  He led the creation of the Digital Archiving Service as part of a wider programme aimed at bringing the department’s digital heap under control, and introducing AI into document review for disposal.  David has spent the past eleven years building the Cabinet Office’s digital records management capability, and this period includes being appointed as the Departmental Records Officer for the Department for Exiting the EU. David strongly believes that a key to solving many of the current issues facing records managers and archivists is to rethink these disciplines through a digital mindset – digital solutions for digital problems. David was awarded MBE for services to innovation in KIM in 2024.    

Dr Kelcey Swain 

Dr. Kelcey Swain is Lead Innovation and AI Architect at the Cabinet Office, specialising in AI governance and digital transformation. With a PhD in Digital Signal Processing, he pioneered the Digital Assurance Navigator to translate ethical policy into machine-readable assurance. Kelcey is driven by the conviction that modern technology can bridge the divide between knowledge and data. He believes that once this gap is closed, AI can be reliably deployed across complex public sector environments. His work focuses on building data-led capabilities that connect complex engineering and operational impact. 

Categories
Uncategorised

LUSTRE/GLOW Workshop Report

On Tuesday 22 July 2025, a one-day workshop titled Government Records and AI was held at the Mercure Piccadilly Hotel, Manchester, as part of the LUSTRE/GLOW initiative led by Professor Lise Jaillant and supported by Loughborough University in collaboration with the UK Cabinet Office. GLOW builds on the work of the LUSTRE project and brings together government professionals, archivists, technologists, and researchers from the UK and US to explore how AI technologies can improve access to government records while addressing issues of sensitivity, metadata, and trust.

The workshop provided a focused platform for participants to discuss current practices, identify shared challenges, and explore responsible, collaborative approaches to using AI across institutional and national contexts.

The presentation slides and abstracts from this workshop can be accessed HERE.

The workshop began with a keynote by Professor Victoria Lemieux of the University of British Columbia, who presented “Clio X: A Web3 Solution for Privacy Preserving AI Access to Digital Archives.” Lemieux introduced Clio X as a decentralised, privacy first Web3 solution that embeds privacy enhancing technologies into archival workflows so institutions can enable AI supported reference and access while maintaining control of sensitive content. Drawing on a user evaluation of a medium fidelity prototype, she discussed adoption barriers around trust, system opacity, costs, and governance. She alsooutlined a participatory, Decentralized Autonomous Organization (DAO) based approach to ethical and transparent uses of AI in cultural heritage and public records contexts.

In the Q& A after the keynote, discussion centred on where sensitivity review should sit and how to keep access workable for researchers, with Professor Victoria Lemieux pointing to privacy-enhancing processing that can reveal useful patterns while withholding identifiers. She also noted the need to run in secure environments with vetted toolsets and controlled access, so archives can support research and appraisal without compromising protection.

The first session Advancing Security and Scalability in Digital Collections, began with Dr James Lappin from the Department for Science, Innovation and Technology, who spoke on “Managing the Digital Heap: Intuitions from Records Management and from Data Science.” that services must scale alongside human-in-the-loop controls and explanations, and that effective practice will draw on both records management and data science to address the full variety, systems, and lifecycles of digital content.

He was followed by Professor Georgina Cosma from Loughborough University, who presented “Preserving and Accessing: Secure AI Solutions for Sensitive Information Systems.” Cosma outlined a Retrieval Augmented Generation framework designed to operate entirely on curated internal sources, avoiding external dependencies. She described how semantic search, classification, and question answering can be delivered in secure environments. She also surveyed agentic RAG, multimodal processing, and the practical challenge of unlearning so that information can be removed from AI systems while preserving transparency and accountability in sensitive settings.

In the second session Evolving User Communities and Institutional Approaches, Dr Adam Nix of the University of Birmingham presented “Understanding Methodological Diversity in Born-Digital User Communities,” mapping four user types, the aggregator, the synthesiser, the fact finder, and the narrator, each with distinct motivations and access needs. Heargued that designing with methodological diversity in mind prevents systems and services from privileging one research approach over others and leads to more equitable, effective access.

Following it, Elizabeth (Lizzie) Hatfield and Andrew Bonnie from the Bodleian Libraries delivered “The Future of AI and Libraries: Opportunities and Challenges at the Bodleian,” situating their exploration of AI in the Bodleian’s mission to enable access for learning and research, outlining the scale, services, and digital strategy that frame the work, and setting out practical challenges they are addressing, including metadata quality, trust, copyright, academic integrity, environmental impact, and infrastructure modernisation as part of a wider digital transformation.

After the break, in the third session Policy Innovation and Practice in Digital Heritage, Dr Katherine McDonough from Lancaster University presented “Historical Maps and the Future of Land Use,” emphasising that historical maps are not mirrors of the past and showing how carefully designed computer vision methods can interrogate very large map collections while remaining sensitive to cartographic genres, an approach that supports policy and land use work without flattening source complexity.

That session continued with Mahin Ali of SVGC presenting “Navigating the Digital Heap with AI,” explaining how machine learning classifiers are being used to accelerate transfer of digital collections to The National Archives while mitigating risk, he outlined a semi-automated triage approach paired with human in the loop sampling and guardrails based on training data sufficiency to increase throughput and deliver public value within agreed risk controls and archival policy.

With those themes in view, the afternoon Plenary Discussion, Collaboration, Access, and Institutional Trust in the Age of AI brought the room together around eight prompts, with each table taking two or three in turn, so the conversation moved from enhancing access and overcoming barriers to ethics, archival principles, future practice, cross sector collaboration, sensitive content detection, and the evolution of work, and it kept circling back to practical workflows, data readiness, transparency, and the need to keep humans in the loop as institutions test and learn.

1. Enhancing Access with AI

Participants tied potential gains in identification and discoverability to very practical steps in existing workflows. They stressed that outputs should be checked for being legally sensible and actionable, then proceed to digital scanning and quality assurance. Any improvement in access should move together with assurances about authenticity, custody, and public confidence, and in some cases, systems may need to maintain or display multiple versions.

2. Overcoming Barriers to AI Adoption

Groups highlighted data readiness and organisational culture as early obstacles. One table cautioned against assuming AI can clean the mess, arguing the foundations must be put in order first or pilots risk seeding mistrust. Concerns also included environmental impact and anxiety about job change among information professionals, with a push to equip staff to use precise technical language rather than umbrella terms. Another table added money, time, confidence, legitimacy, change fatigue, data protection, and the wider political economy, noting that fast tech cultures can collide with archival practice where moving fast and breaking things is not an option.

3. Ethics and Responsibility in AI Use

Participants kept the human central, asked for visible accuracy levels, and called for keeping a record of model changes rather than overwriting them. They discussed bias, hallucinations, and loss of context, including an example where an image classifier began labelling a toy robot as a rock after small changes, which was used to underline the need to educate users about dataset bias. Regulatory frameworks were offered as an analogy for progress with guardrails.

4. Reconciling AI with Archival Principles

The discussion revisited authenticity, provenance, and original order, with suggestions to make the steps that ensure authenticity more explicit and transparent. Several speakers observed that institutions often frame ethical assessment through a risk mitigation lens, which shapes how AI work is evaluated.

5. Looking Ahead: AI and the Future of Archival Practice

Looking five to ten years ahead, groups connected changing record containers and classification practices with risks from synthetic information. One contribution warned about a drift from original meaning if language models repeatedly generate and then consume institutional text, so controls and anchoring in original sources remain important as workflows modernise.

6. Cross-Sector Collaboration for AI Innovation

Collaboration was described as people centred and inevitably slow, because partners work at different speeds and under different incentives. Participants called for shared understanding and shared values across archives, libraries, records management, and academia, and suggested that professional bodies could convene joint efforts instead of working in silos

7. AI for Sensitive Content Detection and Risk Management

Tables distinguished generative systems from other techniques and often preferred trainable classifiers, natural language processing, ontologies, and knowledge graphs for sensitivity review. They stressed that context and risk should guide tool choice and noted that some collections may be too sensitive for AI. Temporal change also matters, for example a classification shift after Brexit, so human-in-the loop supervision remains necessary and prior human decisions can seed training libraries.

8. AI and the Evolution of Archival Work

Participants anticipated shifts in competencies toward clearer technical vocabulary and better judgement about when a particular tool is fit for purpose. They discussed translating from paper-based sensitivity review to digital sensitivity review by building on existing frameworks rather than reinventing processes, and they observed that information is increasingly stored around people rather than in rigid file plans, with implications for future appraisal and description.

After the breakout group session, in Keynote and Closing Remarks, David Canning MBE of the Cabinet Office presented “Everything, Everywhere, All at Once,” reflecting on the constant challenges that archivists and records managers face in keeping up with and adapting to new and emergent technologies. He outlined how we can keep focused through the hype, and some of the steps already being taken by the United Kingdom Government to protect and preserve the official record.

The workshop concluded with a Q&A following David Canning’s keynote, drawing together the day’s discussion. Topics included the shift from paper to fully digital workflows and an unbroken chain of information, how to treat Teams chats, video meetings and consumer messaging alongside email as formal records, and the need for clear retention and privacy rules to guide practice.

Overall, the GLOW workshop offered practical insights into privacy by design, secure and scalable access, user-centered discovery, and the evolution of record keeping, and it highlighted the importance of continued cross-sector collaboration to turn these conversations into shared guidance and pilots.

Categories
Uncategorised

Historical Maps and the Future of Land Use 

22nd July 12:10 – 12:25

Speaker: Dr Katherine McDonough

Abstract: Historical maps – from planning documents to national cadastres – are increasingly being scanned and shared as digital image collections. It is tempting to treat all representations of past landscapes as truthful reflections of built and natural environments, but experts in historical cartographic genres know that maps are not mirrors, nor do they follow the same rules as more recent earth observation imagery. In this talk, I demonstrate how the careful application of computer vision tools can help researchers, policy makers and land managers interrogate big, historical map data while remaining sensitive to the complexities of these sources. 

Bio: Katherine McDonough is a Lecturer in Digital Humanities in the Department of History at Lancaster University. She is a historian of eighteenth-century France and a specialist in the spatial digital humanities. She founded the MapReader software library with colleagues at The Alan Turing Institute and continues to lead this initiative to work with historical maps as data. 

Categories
Uncategorised

Data Protection and AI: how to apply the data protection principles to the use of information in AI systems.

Speaker: Alister Pearson, Principal Policy Advisor for AI and Data Science, Information Commissioner’s Office (ICO).

Abstract: Artificial Intelligence (AI) has the potential to bring significant benefits for researchers. LUSTRE has identified that AI can be used to identify sensitive materials in a mass of born-digital records to make non-sensitive materials accessible and can also serve to search vast amounts of data when keyword searches would not be effective. However, the use of AI can also introduce or exacerbate risks to people’s right to privacy as well as other rights that relate to the processing of their personal information. In this presentation, I will discuss some of the main data protection risks of using AI in a research context, as well as ways to mitigate these risks. This will include considering risks during both the development and deployment of AI. My presentation will also include the ways that the Information Commissioner’s Office can further help researchers via the different services that we offer, including our recently launched Innovation Advice service. One of the aims of the presentation is to illustrate that data protection provides a framework to process people’s personal data, rather than a barrier

Bio: Alister Pearson is a Principal Policy Advisor for the AI and Data Science Team at the Information Commissioner’s Office (ICO). He was part of the team that produced the ICO’s Guidance on AI and Data Protection and led on the development of the AI and data protection risk toolkit, which won the Accountability prize at the Global Privacy Assembly Awards 2022.

Categories
Uncategorised

Using AI to improve Excel skills

Speaker: Angie Campbell, Preservation Manager, Public Record Office of Northern Ireland (PRONI).

Abstract: Using AI (Chat GPT) as a tool to assist with Excel problems and, hopefully, improve skills. Engaging in a conversation with Chat GPT, non-technical individuals can use step-by-step guidance on troubleshooting Excel issues, learn new features and functions, and gain practical knowledge to enhance their Excel proficiency.

Bio: Angie Campbell is a Preservation Manager within PRONI. Angie has over 12 years experience in collections management, including preservation, storage, and oversight of the document production service. Angie’s team includes a mix of professional/curatorial, administrative and support grade staff. Her current role involves the assessment of collections held in out-storage.

Categories
Uncategorised

I, Historian: Researching the past in the age of Artificial Intelligence

Speaker: Dr David Brown, Senior Researcher at the Virtual Treasury of Ireland

Abstract: Although Artificial Intelligence (AI) burst into the public consciousness towards the end of 2022 with the launch of ChatGPT, a chatbot made by a company with close links to Microsoft, AI technology has been a core technology for the Virtual Treasury of Ireland since the early days of the project. Starting in 2018, the Virtual Treasury of Ireland has developed a suite of deep-learning ‘models’, perfectly curated transcriptions, to train an AI system to read digital images of historical sources relating to Ireland and convert these into searchable text files. The automatic conversion of handwritten historical documents into searchable text is the latest major step in the digitisation of our written cultural heritage. Digital images can be expensive to create, costly to store, and can degrade over time. Moreover, as digital images are so easy to produce, they have a habit of proliferating yet are no more searchable than the original records they were intended to be surrogates for. AI enables us to preserve all of the benefits of digital images while eliminating these drawbacks. The further digitisation of the images into text makes them searchable, portable and easy to store. This new ability to find a person, place or event among thousands or millions of pages of handwritten documents is the first step in an exciting AI-enabled world that will enable ever more complex historical hypotheses to be tested and questions to be answered. Emerging technologies such as the Large Language Models (LLMs), upon which ‘human-seeming’ interfaces such as ChatGPT are based, promise to enable ever more complex research questions at the touch of a button. Current research at the VRTI is demonstrating that LLMs can improve the accuracy of our transcriptions, summarize the most important information contained in dense and lengthy documents, extract people and places from text to incorporate into our Knowledge Graph and even sort documents into chronological order.

Bio: Dr David Brown is Senior Researcher at the Virtual Treasury of Ireland, a project based at Trinity College Dublin that aims to recreate a digital model of the Public Record Office of Ireland with its contents, destroyed in 1922. David has published widely on early modern empires with a particular focus on trade and finance and has been involved in digitisation projects, in various guises, for over 30 years.

Categories
Uncategorised

RAISE: Responsible AI iS an Enabler 


Speaker: Michaela Black, Professor of Artificial Intelligence at Ulster University

Abstract: The presentation will look at what is Responsible AI is and how it can help build trust in adoption of AI.

By doing so more data providers will be keen to contribute their data and thus more benefits can be grown from the learning from AI.

Speaker’s bio: Michaela Black, Professor of Artificial Intelligence at Ulster University, her career has seen her apply Artificial Intelligence and gamification to a wide range of domains including: telecoms, healthcare, finance, education and marketing with 60+ publications and securing £7M+ in research and pedagogic funding. She has actively engaged in a wide range of AI projects securing funding from a wide range of funders including coordinating an EU-funded project MIDAS (Meaningful Integration of Data, Analytics and Services) 2016-2020. This 40-month project, secured €4.5m in funding from the European Union’s Horizon 2020 programme. It delivered Data Science hardware and software solutions to connect fragmented health data and enable policymakers across 4+ European countries to analysis data and enhance policies across a range of diverse topics such as: mental health, obesity, diabetes and looked after children, as well as sharing and ensuring uptake of excellent data access practices such as MIDAS Honest Broker Service. She is also a member of a consortium who has just secured European funding for LUCIA: Understanding Lung Cancer related risk factors and their Impact launched January 2023. Michaela is a current co-director of the engage technology, a stakeholder participation system, proven and effective platform for gaining participant input and key consensus on relevant topics using real-time digital technologies. engage has delivered successful engagements in partnership with organisations across government, industry, professional bodies, PPIs, NGOs and academia, on the island of Ireland and across Europe. This work has influenced topics such as Brexit, work force planning in Department of Health and Matrix NI Policy on Women in STEM. An active STEM ambassador, Michaela strives to encourage more women to join this field. Actively part of the Athena Swan (AS) team at UU to promote diversity and equality. She secured EPSRC funding with University of Bath Inclusion Matters Scheme, to host a datathon which aimed to improve equality, diversity and inclusion within the engineering and physical sciences.

Categories
Uncategorised

Two alternative pathways for the application of AI for recordkeeping purposes in live systems


Speaker: James Lappin

Abstract: Originating organisations have digital material at all stages of the records lifecycle: in live systems, in legacy systems, and awaiting appraisal for possible transfer to a historical archive.  Organisations will need help from AI at each stage, but they will face different challenges at each stages in developing AI models; in deploying AI models; and/or in basing decisions and actions on AI models. 

This talk will look at the application of AI models in live systems such as the Microsoft 365 cloud suite, where end-users are working and adding content into a wide variety of aggregations (email accounts, SharePoint sites, OneDrive accounts, Teams, Chat accounts etc.).  It will look at the special challenges and opportunities the deployment of AI in live systems poses. 

Two alternative pathways for the application of AI models in live systems will be compared and contrasted:

  • One pathway uses the existing way that content is aggregated as its starting point, working within existing aggregations to make them more precise, useful and manageable; 
  • The other pathway works across the entirety of an organisation’s ‘digital heap’ (or over the entirety of the organisations M365 implementation) to identify and label important content.

The talk uses both recordkeeping theory, and experience with digital records over the past thirty years, to arrive at some predictions as to which of the two pathway is likely to provide the safer and more predictable route to using AI to achieve improvements in recordkeeping .


Speaker Bio
: James Lappin has worked in the field of archives and records management for thirty years as a practitioner, consultant, researcher, policy advisor, presenter, blogger, podcaster and cartoonist.