Agentic systems | Coding-agent evaluation | Responsible AI | Multilingual NLP

Oyinkansola Onwuchekwa

AI Researcher and Engineer | England, United Kingdom

I build and evaluate agentic systems, with a focus on coding-agent controls and trace-level evaluation. My wider research covers responsible AI, multilingual NLP, low-resource African languages, and culturally grounded emotion modelling.

Current programme Agent traces to controlled tool use See the route
0.5.0 Secure Agent Gateway release
320 + 160 Core and adversarial control-bench traces
18,402 Held-out AfriSenti test items
2 Merged upstream pull requests
Oyinkansola Onwuchekwa at the University of Hull

Public projects

Source code, release records, benchmark reports, and research notes are linked with each project. Reported figures come from held-out evaluations or versioned release records.

Research programme

From trace evidence to controlled action

I study what an agent did, how a monitor can identify the point of failure, and which controls should sit before tool execution.

  1. 01TraceRecord the agent’s actions
  2. 02MonitorLocate policy-relevant evidence
  3. 03ControlCheck the request before execution
  4. 04EvaluateMeasure safety and task utility

Monitor benchmark | Version 0.4.1

Agentic Security Control Bench

A contrastive benchmark for testing whether coding-agent monitors ground control decisions in policy-relevant trace evidence. The release contains 320 core traces and a separate 160-trace adversarial holdout.

Evaluation measures unsafe-action prevention, permitted-task retention, evidence grounding, intervention timing, calibration, approval burden, and paired bootstrap intervals.

320 core traces160 adversarial tracesDOI archived

Python package | Version 0.3.0

Low-Resource NLP Toolkit

African-language text normalisation, selective language routing, code-switch audits, emotion-label mapping, and coverage-aware evaluation. The package runs without an API key or model download.

On 18,402 held-out AfriSenti test tweets, the development-selected rejection policy reached 89.65% accuracy on accepted items at 74.03% coverage.

PyPI packageOffline evaluationDocumentation site

Dataset | Version 0.1.0

Coding Agent Failure Atlas

A reproducible dataset of 120 synthetic coding-agent traces across 12 safety and security failure modes, including integrity problems. Each trace links the failure evidence to its first intervention point and a safer counterfactual.

Evaluation harness

Coding Agent Monitor Lab

A Python harness for scanning agent traces and evaluating deterministic monitors. Findings include localised evidence and a remediation field, with JSON reports for regression testing or later model comparison.

Multilingual benchmark

Multilingual DimStance Baselines

CPU-friendly valence-arousal baselines for German, English, Nigerian Pidgin, Swahili, and Chinese. Evaluation uses fixed training budgets and grouped resampling, with source-record bootstrap intervals.

Aspect conditioning reduced test macro RMSE from 1.343 to 1.323. The language-adaptive result is reported separately by language.

Accepted upstream contributions

These fixes were reviewed and merged by maintainers of established Python AI projects. The pull request pages retain the review discussion and CI record.

Haystack Merged 15 June 2026

Reject non-list QueryExpander replies

Fixed a case where a valid JSON object containing a string produced one query per character. The change added type validation and regression coverage; its release note records the behaviour.

View merged pull request
Sentence Transformers Merged 6 July 2026

Warn on non-binary contrastive labels

Added a one-time warning for labels outside the documented binary form, while preserving existing weighted-loss workflows. Focused tests cover ordinary and non-binary inputs without downloading model weights.

View merged pull request

Public research outputs

Published records and versioned technical outputs with permanent public links.

2024

Peer-reviewed paper

BDA at SemEval-2024 Task 4: Detection of Persuasion in Memes Across Languages with Ensemble Learning and External Knowledge

Proceedings of the 18th International Workshop on Semantic Evaluation, pages 123-132.

ACL Anthology record
2025

Conference abstract and presentation

Locally-responsible Artificial Intelligence frameworks

Co-authored work from the DAIL-ICH project, presented at Digital Humanities 2025 in Lisbon. The full title and author record appear in the published book of abstracts.

DH2025 Book of Abstracts

Research and engineering practice

Research

I am completing a PhD in Data Science and Artificial Intelligence at the University of Hull. My research examines multilingual emotion classification in code-mixed Afrobeats lyrics, with work on retrieval systems and label provenance.

Engineering

My recent engineering work covers agent tool policy, coding-agent monitor evaluation, research automation, and reproducible NLP systems.

Service and public engagement

Academic and professional service

  • Member, UKRI EPSRC and NERC Peer Review Colleges
  • Reviewer, International Conference on Learning Representations
  • Reviewer, Deep Learning Indaba
  • Associate Fellow of Advance HE (AFHEA)
  • Professional Member, BCS, The Chartered Institute for IT

Speaking and community work

  • Panel speaker, 6th European Chatbot and Conversational AI Summit
  • Co-author and presenter, Digital Humanities 2025, Lisbon
  • Invited speaker, PyCon Lithuania 2025
  • Lead organiser and speaker, Towards Transparent and Responsible AI Conference
  • Invited speaker, Pint of Science UK, Hull
  • Featured contributor, BBC News interview for National AI Day

My community work includes academic support with IntoUniversity Hull East, research mentoring through the University of Hull and Nuffield Research Placements, Tech Help Group, and the Hull Afro-Caribbean Association.

Python PyTorch Hugging Face scikit-learn RAG Coding-agent evaluation Policy testing Responsible AI GitHub Actions Docker

Research and engineering enquiries

For research collaboration, technical discussion, speaking, or professional enquiries, use the links here.