We study how AI systems fail people, and build the evidence to fix them.
SRLab is a registered not-for-profit research laboratory based in Canada. We test language, vision, and agentic AI systems for fairness, safety, and accountability, and we publish what we find: benchmarks, datasets, software, and peer-reviewed research that others can inspect, reuse, and challenge.
- Fairness and biasBenchmarks and detectors for text, image, audio, and video models, reported by demographic group.
- Agent evaluation and governanceTrajectory-level evaluation, transparency records, and trust, risk, and security management for AI agents.
- Synthetic media and deceptionDisinformation and deepfake detection, and models that resist repeating known falsehoods.
- Environmental cost of AICarbon, energy, and water accounting across model lineages, with public dashboards.
- Legal name
- Scientific Research-AI Lab, operating as SRLab
- Legal form
- Not-for-profit corporation, Canada Not-for-profit Corporations Act, 2025
- Corporation number
- 1697413-9
- Contact
- contact@srlab.ai
About the lab
SRLab was founded in 2025 and incorporated under the Canada Not-for-profit Corporations Act. The research programme is led by the lab's Research Director and carried out with a small group of scientists, graduate students, and interns from partner institutions in Canada, the United States, and Europe.
Our earliest work was on bias in news text, published at AAAI and ACM AIES and released as open-source software that is still in use. From there the lab grew into multimodal fairness benchmarks, safety evaluation for language models, and, most recently, the governance and evaluation of AI agents. Our work is supported through the EU Horizon Europe programme and through collaborations with Canada's principal AI and health research bodies.
Recent
-
September 2026 Keynote at the GREEN-AI Workshop at ECML-PKDD 2026 in Naples on Data and Impact Accounting, our ICML 2026 spotlight position paper.
-
September 2026 Proposed a half-day workshop, The Hidden Compute Cost of AI Safety Evaluation, for IASEAI 2027, building on our Data and Impact Accounting work.
-
June 2026 unbias-plus demonstrated to the Prime Minister of Canada during a visit to the Vector Institute.
-
June 2026 Presented HumaniBench and SONIC-O1 in the UNLOCK benchmarking session at HAICON 2026, Helmholtz Munich.
-
January 2026 F-DPO accepted at ACL 2026 Findings; SONIC-O1 released; VLDBench published in Information Fusion; TRiSM for agentic AI published in AI Open.
-
January 2026 AI governance panel at AAAI 2026 in Singapore alongside researchers from IBM, Google DeepMind, and Microsoft.
-
October 2025 Six acceptances at NeurIPS 2025 across fairness, sustainability, and agent safety; keynote at IEEE SMC 2025.
Recognition
-
Responsible AI Leader of the Year, Women in AI Summit and Awards 2025 (North America).
-
The Peak's Emerging Leaders 2026 in artificial intelligence (Canada).
-
Stanford/Elsevier list of the world's top 2% most-cited scientists, 2024 and 2025.
-
CIHR Health System Impact Fellowship and AI4PH award, for responsible AI in health equity and public health.
-
Co-applicant on AIXPERT, one of three projects selected from 137 proposals under Horizon Europe's trustworthy AI call.
-
Editorial and review service: reviewer for Nature and Nature Machine Intelligence; editorial board of Springer Discover Computing; NeurIPS main-track programme committee; CIHR and Killam peer-review committees.
How we work
- We release what we buildBenchmarks, datasets, and code are published under open licences so that results can be reproduced and criticised.
- We keep clear boundariesMembers hold positions at other institutions; work done there is attributed there, and lab resources are not mixed with external projects.
- We prefer measured claimsMuch of our research asks whether AI evaluations mean what they appear to mean. We hold our own findings to the same standard.
Where we would like to go
Over the next few years we want to grow from a volunteer-led lab into a stable Canadian research organisation with funded positions for early-career researchers, particularly on the evaluation science of agentic AI, synthetic-media forensics, and the environmental cost of AI systems. We are seeking public, philanthropic, and industry partners who share that aim.
Tools and proofs of concept
Software, benchmarks, and demonstrators we build and maintain. Everything here is public and free for research use.
-
unbias-plusTool
Bias detection and rewriting toolkit built on a fine-tuned Qwen3-4B. Classifies bias type and severity, locates biased spans, and returns a neutral rewrite in one structured output. Demonstrated to the Prime Minister of Canada in June 2026.
-
FairSense-AgentiXProof of concept
Agentic fairness and risk analysis platform: ReAct reasoning loops, dynamic tool selection, and self-critique over text, images, and datasets. FastAPI backend with streaming and a React front end.
-
FairSense-AITool
Toolkit for bias auditing, emissions reporting, and explainability across large-scale AI systems, with responsible-AI compliance checks.
-
SONIC-O1Benchmark
Human-verified benchmark for audio-video understanding: 231 videos, about 60 hours, 4,958 questions across 13 domains, with results reported by demographic group. Public leaderboard.
-
HumaniBenchBenchmark
Human-centred benchmark for large multimodal models organised around seven principles, including fairness, empathy, robustness, and multilingual ability. Presented at HAICON 2026.
-
VLDBenchBenchmark
Benchmark for multimodal disinformation detection with regulatory alignment. Published in Information Fusion (2026).
-
Data and Impact Accounting (DIA)Tool
Tracks the cumulative carbon and water footprint of open-source models and their derivatives, with a public dashboard. ICML 2026 spotlight position paper; keynote at GREEN-AI, ECML-PKDD 2026.
-
Educational visualisersProof of concept
Interactive teaching tools accepted to the NeurIPS 2025 Educational Program: the climate cost of AI expressed in human-scale equivalents, and model immunisation, prompt-injection defence, and red-teaming explained through the vaccine analogy.
-
GreenBenchBenchmark
Benchmark for the energy footprint of AI models across training and inference settings.
-
Agentic transparencyFramework
Survey and framework for recording what an agent did and why, with the Minimal Explanation Packet as an audit-ready record at the end of each task.
-
Responsible agentic reasoning (R²A²)Proof of concept
Agents that carry fairness, privacy, and auditability checks through each reasoning step rather than only at the final output.
-
Dbias, NBias, and GUS-NetTool
Earlier bias-detection releases: a configurable pipeline for news text (Dbias), token-level bias entities (NBias), and span-level generalisation, unfairness, and stereotype tagging (GUS-Net).
Consulting
We work with companies, public bodies, and other research organisations that need an independent assessment of how their AI systems behave. Engagements are scoped in writing, staffed by lab scientists, and priced to cover costs; any surplus funds the lab's research programme and student placements.
-
Independent AI audits
Evaluation of a deployed or pre-deployment AI system against fairness, safety, robustness, and factuality criteria, using our own benchmarks and tooling. You receive a written report with reproducible evidence, not a certificate.
-
Bias and fairness assessment
Measurement of demographic and linguistic bias in language, vision, and multimodal models, including hiring, content moderation, health, and news applications. Results reported by group, with recommended mitigations.
-
Agentic AI governance
Design review of agent systems: trajectory logging, tool-use controls, transparency records, and trust, risk, and security management, drawing on our TRiSM and Minimal Explanation Packet frameworks.
-
Environmental footprint accounting
Carbon, energy, and water accounting for training, fine-tuning, and inference workloads, aligned with emerging EU AI Act reporting expectations, using Data and Impact Accounting.
-
Regulatory and policy readiness
Gap analysis against the EU AI Act, Canadian AI guidance, and sector standards, with a prioritised plan for documentation, testing, and disclosure.
-
Training and workshops
Half-day to multi-day sessions for engineering, product, policy, and executive teams on responsible AI evaluation, bias detection, agent safety, and sustainable AI practice.
How an engagement works
- ScopeA short call, then a written scope with questions, systems in view, data access, timeline, and cost. Typical engagements run from a two-week assessment to a multi-month programme.
- AssessLab scientists run the evaluation using our benchmarks and tooling, with interim findings shared as they emerge.
- ReportA written report with reproducible evidence and prioritised recommendations, plus a walkthrough for your team. We do not accept work that requires us to withhold findings from the people affected by a system.
Papers
Our work falls into four themes: fairness and bias in text and multimodal models; evaluation and governance of AI agents; synthetic media and deception; and the environmental cost of AI. A selection is below; work under anonymous review is not listed until it is accepted. Full lists are on Google Scholar under each author.
Agents: evaluation and governance
- Evaluating and regulating agentic AI: a study of benchmarks, metrics, and regulation. Farooq, Raza et al., Emmanouilidis. Information Fusion, 2026.
- TRiSM for agentic AI: a review of trust, risk, and security management in LLM-based agentic multi-agent systems. Raza, Sapkota, Karkee, Emmanouilidis. AI Open, 2026.
- Transparency in agentic AI: a survey of interpretability, explainability, and governance. Raza, Radwan, Chaduvula, Alinoori, Emmanouilidis. Preprint, 2026.
- From features to actions: explainability in traditional and agentic AI systems. Chaduvula, Ho, Kim, Narayanan, Alinoori, Garg, Ramachandram, Raza. Preprint, 2026.
- Who is responsible? The data, models, users or regulations? A comprehensive survey on responsible generative AI for a sustainable future. Raza, Qureshi, Zahid et al. ACM Computing Surveys, 2025.
- Developing safe and responsible large language models. Raza, Bashir et al. Machine Learning, 2025.
Fairness and bias
- FairLens: benchmarking fairness in vision-language models for high-stakes decision-making. Raza et al. Preprint, 2026.
- Bias in the picture: benchmarking VLMs with social-cue news images and LLM-as-judge assessment. Narayanan, Khazaie, Raza. NeurIPS 2025 workshop (Evaluating the Evolving LLM Lifecycle).
- ViLBias: detecting and reasoning about bias in multimodal content. Raza et al. Preprint, 2025.
- SONIC-O1: a real-world benchmark for evaluating multimodal LLMs on audio-video understanding. Radwan, Emmanouilidis, Tabassum, Pandya, Raza. 2026.
- HumaniBench: a human-centric framework for large multimodal models evaluation. Raza et al. NeurIPS 2025 poster; under revision for ACM TIST.
- Beyond content: how grammatical gender shapes visual representation in text-to-image models. Raza et al. EMNLP 2025 Findings.
- LinguaMark: do multimodal models speak fairly? A benchmark-based evaluation. Raval, Narayanan, Khazaie, Raza. ASONAM 2025.
- RESPECT: a framework for promoting inclusive and respectful conversations in online communications. Raza et al. Natural Language Processing Journal, 2025.
- MBIAS: mitigating bias in large language models while retaining context. Raza, Raval, Chatrath. ACL 2024 WASSA workshop.
- NBIAS: a natural language processing framework for bias identification in text. Raza et al. Expert Systems with Applications, 2023.
Synthetic media and deception
- VLDBench: evaluating multimodal disinformation with regulatory alignment. Raza, Bashir, Emmanouilidis, Shah et al. Information Fusion, 2026.
- F-DPO: reducing hallucinations in LLMs via factuality-aware preference learning. Raza et al. ACL 2026 Findings.
- Just as humans need vaccines, so do models: model immunization to combat falsehoods. Raza et al. IJCNN / IEEE WCCI, 2026.
- Fake news detection: comparative evaluation of BERT-like models and large language models with generative AI-annotated data. Raza, Paulen-Patterson, Ding. Knowledge and Information Systems, 2025.
- FakeWatch: a framework for detecting fake news to ensure credible elections. Raza et al. Social Network Analysis and Mining, 2024.
- Fake news detection based on news content and social contexts: a transformer-based approach. Raza, Ding. International Journal of Data Science and Analytics, 2022.
Environmental cost of AI
- Sustainable open-source AI requires tracking the cumulative footprint of derivatives. Raza et al. ICML 2026, spotlight position paper.
- Optimizing large language models: metrics, energy efficiency, and case study insights. Khan, Motie, Kocak, Raza. Preprint, 2025.
- FairSense-AI: responsible AI meets sustainability. Raza, Chettiar, Yousefabadi, Khan, Lotif. Preprint, 2025.
Recommender systems and applied work
- Review-based recommender systems: a survey of approaches, challenges and future perspectives. Hasan, Rahman, Ding, Huang, Raza. ACM Computing Surveys, 2026.
- A comprehensive review of recommender systems: transitioning from theory to practice. Raza et al. Computer Science Review, 2026.
- Progress in context-aware recommender systems: an overview. Raza, Ding. Computer Science Review, 2019.
- Large-scale application of named entity recognition to biomedicine and epidemiology. Raza, Reji, Shajan, Bashir. PLOS Digital Health, 2022.
- Clinical Application of Detecting COVID-19 Risks: A Natural Language Processing Approach. Bashir, Raza, Kocaman, Qamar. Viruses, 2022.
- A jamming attack detection technique for opportunistic networks. Singh, Woungang, Dhurandher, Khalid. Internet of Things, 2022.
- Reinforcement learning-based fuzzy geocast routing protocol for opportunistic networks. Khalid, Woungang, Dhurandher, Singh. Internet of Things, 2021.
In progress
-
Reproducibility of bias-detection tools
An audit of whether span-level bias detectors give the same verdict when the sampling seed, inference platform, or surrounding text changes.
-
Harness-aware evaluation of agents
A survey arguing that an agent's benchmark score belongs to the whole evaluation set-up (model, harness, environment, and evaluator), not the model alone.
-
Deepfake forensics
Audio-visual deepfake detection that generalises across generators: a survey, a trace-labelled forensic benchmark, and detection models that route on signal reliability.
-
Does cheaper evaluation change the conclusion?
Whether quantisation, batching, and benchmark reduction preserve conclusions about accuracy, fairness, and bias, and what each saves in energy and water.
Talks
Keynotes, panels, and invited talks by lab members. Invitations can be sent to contact@srlab.ai.
-
Sep 2026
GREEN-AI Workshop, ECML-PKDD 2026, Naples
Sustainable open-source AI requires tracking the cumulative footprint of derivatives keynote -
Jun 2026
HAICON 2026, Helmholtz Munich
Current status of the benchmarking field: lessons learned (UNLOCK initiative) -
Jan 2026
AI Governance Workshop, AAAI 2026, Singapore
Panel on AI governance, safety, and responsible AI with IBM Research, Google DeepMind, and Microsoft Research - Oct 2025
- Oct 2025
-
Jun 2025
ResearchTrend.AI
HumaniBench: a human-centric framework for large multimodal models evaluation -
Nov 2024
Conestoga College AI Symposium
AI, data privacy, and ethics -
Jul 2024
Cohere
Mitigating biases in text and LLMs: challenges and strategies -
Mar 2024
ECIR 2024, Glasgow
Dissecting news narratives and misinformation -
Mar 2024
Lassonde School, York University
Enhancing fairness in large language models -
Jan 2024
Rogers Catalyst, Toronto Metropolitan University
Cybersecurity and AI: who loses? The case for more inclusive engagement -
Aug 2023
Human-Medical AI Symposium, AAAI 2023
Fairness in machine learning meets equity in healthcare -
Jun 2023
ICA 2023
Accuracy meets diversity in a news recommender system keynote -
Apr 2023
Rice University
Connecting fairness in machine learning with public health equity
People
Board of directors
-
Dr. Syed Raza Bashir
Founder, Chair of the BoardFounder and Chair of the Board of SRLab.ai, with more than 25 years of experience across applied AI, software engineering, project and program management, higher education, research, and technology leadership.
More
Holds a PhD in Computer Science from Toronto Metropolitan University and is a PMP-certified professional with expertise in machine learning, responsible AI, software engineering, recommender systems, biomedical informatics, and fairness and bias in AI. Has led multidisciplinary technical projects, managed end-to-end research and software initiatives, and mentored students and engineering teams across academic, industry, and publishing environments, including as Scientific Managing Editor in Computer Science at Elsevier. Visiting faculty and capstone supervisor at Sheridan and Humber colleges. Research record includes publications in multimodal disinformation, responsible large language models, AI fairness, privacy, healthcare AI, and recommender systems. Provides strategic and governance leadership, helping shape the lab's research direction, partnerships, and mission to advance impactful, responsible, and applied AI research.
-
Dr. Shaina Raza
Chief ScientistApplied Machine Learning Scientist in Responsible AI at the Vector Institute and co-lead of Work Package 3 of the Horizon Europe AIXPERT project. Former CIHR Health System Impact Fellow working on equity in public-health AI. Responsible AI Leader of the Year (Women in AI, North America, 2025), one of The Peak's Emerging Leaders 2026, and on the Stanford/Elsevier top 2% list. Sits on CIHR and Killam peer-review committees. Principal investigator on the lab's research programme.
-
Dr. Khurram Khalid
Director and ScientistResearcher in opportunistic networks, secure routing, and energy-efficient protocols, with publications in Internet of Things (Elsevier) and IEEE venues. Leads the lab's work on secure and resilient systems.
-
Shujaat Feroze
Director, IT InfrastructureM.S., University of Wollongong. More than a decade of enterprise IT leadership at Telstra, Qantas, Saunders International, and Comcare. Responsible for the lab's cloud, security, and operational infrastructure.
Scientific advisory board
-
Dr. Rizwan Qureshi
Scientific AdvisorSenior Research Fellow at Massachusetts General Hospital, Harvard Medical School, where he works on foundation models for women's health and sex-specific disease. Previously at MD Anderson Cancer Center and the University of Central Florida.
More
Senior member of IEEE and DAAD AI Fellow; about 6,000 citations. Co-author of the lab's ACM Computing Surveys review of responsible generative AI.
-
Prof. Athanasios Vasilakos
Scientific AdvisorDistinguished Professor at the Center for AI Research, University of Agder, Norway, and Dean of Class VI of the European Academy of Sciences and Arts.
More
Web of Science Highly Cited Researcher with more than 80,000 citations across AI, cybersecurity, and the Internet of Things. Advises on European partnerships and research direction.
Management and operations
-
Imran Liaquat
Advisor, Strategy and OperationsManagement consultant with more than two decades of international experience helping organisations develop and execute strategy, transform operations, and deliver complex business initiatives. A PMI Authorized Trainer, he has trained and mentored thousands of professionals across North America, Europe, and the Middle East.
More
His expertise spans strategy development and execution, the Balanced Scorecard, and KPI and OKR frameworks, together with portfolio, programme, and project management, risk and quality management, and organisational change. At SRLab he advises on organisational strategy, governance, operational effectiveness, and funding decisions, helping align the research portfolio with sustainable growth, strategic priorities, and partnerships.
Students and interns
Graduate students and research interns join the lab for defined projects, usually leading to a co-authored publication or a public release.
-
Mahveen Raza
Research InternWorks on AI sustainability, bias in generative models, and deepfake detection.
More
Co-author of four NeurIPS 2025 workshop papers (GenProCC and Muslims in ML), including studies of carbon literacy for generative AI and occupational stereotypes in text-to-image models, and lead author of the lab's NeurIPS 2025 educational materials, which she presented in San Diego. Currently working on the deepfake forensics programme. Portfolio
Women and communities left out of AI
AI systems are trained on the world as it has been recorded, and the record is uneven. Women, racialised people, newcomers, and people outside English-speaking, high-income settings are the groups most often misdescribed or overlooked by these systems, and least often in the room when they are designed. A large part of our research exists to make that visible and measurable.
What our research has found
Image generators change who they draw depending on the grammatical gender of a word, and assign occupations along stereotyped lines. Vision-language models rate the same content differently depending on the apparent gender or ethnicity of a face. Audio-video models perform unevenly across the demographic groups in our SONIC-O1 benchmark. Bias-detection tools themselves give different verdicts depending on how they are run. These are not abstract concerns; they shape hiring screens, content moderation, health information, and news feeds.
How the lab is built
SRLab's research programme is led by a woman. Our board includes a researcher whose full-time work is AI for women's health. Our founders are immigrants to Canada who built their careers here, and most of our students and interns come from groups that remain underrepresented in AI research. We do not think that makes our work better by itself, but it does mean the questions we ask start from lived experience rather than from a checklist.
What we are committing to
Every benchmark we release reports results by demographic group, not only in aggregate. We give priority in student and intern placements to women and to people from underrepresented communities, and we pay for the work wherever funding allows. We publish our methods so that community organisations, regulators, and journalists can check AI systems themselves rather than take a vendor's word. And we are seeking partners, including through Women and Gender Equality Canada and provincial programmes, to turn these findings into changes in how institutions procure and audit AI.
Partners and affiliations
Institutions our members belong to or collaborate with. Listing here does not imply endorsement by the institution.
Universities and research institutes
- Vector Institute for AI
- University of Toronto
- Toronto Metropolitan University
- University of Agder
- Harvard Medical School / MGH
- Cornell University
- University of Groningen
- University of Central Florida
- University of Guelph
- University of Tennessee HSC
- University of Wollongong
- Clarkson University
- Arizona State University
- Mayo Clinic
- Helmholtz Munich
- Sheridan, Humber, and Conestoga Colleges
Programmes and funders
- Horizon Europe (AIXPERT)
- Canadian Institutes of Health Research
- CIFAR AI and Society
- NSERC
- Partnership on AI
Professional bodies
- IEEE
- ACM
- European Academy of Sciences and Arts
- NeurIPS, ICML, AAAI, ACL
- IASEAI
Governance and research integrity
Scientific Research-AI Lab, operating as SRLab, is incorporated under the Canada Not-for-profit Corporations Act (corporation number 1697413-9, incorporated 7 May 2025) and is governed by its board of directors. The board approves the research programme, budget, and any partnership or funding agreement.
Disclosure of affiliations
Every member's external affiliation, academic or industrial, is listed on this page. Work carried out at another institution is attributed to that institution in publications and on this site.
Separation of resources
Lab funds, data, and computing are used only for lab projects. Members do not use lab resources for work belonging to their employers, and do not bring employer resources into lab projects without a written agreement.
Conflicts of interest
Before a grant application or partnership is signed, the board reviews potential conflicts. A director with a conflict does not take part in the decision. Conflict declarations are recorded in the board minutes.
Open outputs
Publications are made available as open-access preprints. Datasets and code are released under permissive licences unless a data-provider agreement prevents it, in which case we say so.
Questions about our governance can be sent to contact@srlab.ai.
Working with us
We are looking for funding partners, consulting clients, host institutions for student placements, and collaborators who want to test AI systems against real-world harms rather than benchmark leaderboards.
Our current priorities are the evaluation science of AI agents, forensic detection of synthetic media, and accounting for the environmental cost of AI. We are pursuing support through Canadian programmes such as NSERC, CIFAR, and Innovation, Science and Economic Development Canada, through Women and Gender Equality Canada's leadership and economic opportunity programmes, through follow-on Horizon Europe calls, and through foundations such as Schmidt Sciences and Mozilla.
If you would like to talk, write to us. We reply to every message.
Tell us about the system you want tested, or the programme you want to fund.
contact@srlab.ai