Zeinab Zolaktaf

Zeinab Zolaktaf

Vancouver, Canada

I'm an applied scientist building generative AI systems, from LLM evaluation and experimentation tooling to production applications, often across teams. Lately I've been using agentic AI to move much faster than I ever could on my own.


Experience

Amazon - Applied Scientist (2022–Present), Vancouver, Canada

Elpha Secure - Lead Applied Scientist (2021–2022), Vancouver, Canada

Olyns - Machine Learning Consultant (2021), San Francisco, USA

Georgian - Applied Data Scientist (2020–2021), Toronto, Canada

EhsAI - NLP and Machine Learning Scientist (2019–2020), Vancouver, Canada


Education

PhD, University of British Columbia (2012–2019), Vancouver, Canada

MSc, Dalhousie University (2009–2012), Halifax, Canada

BSc in Computer Software Engineering, Isfahan University (2004–2008), Isfahan, Iran


Publications

Workload-Aware Query Recommendation Using Deep Learning (EDBT, 2023)

This paper studies how workload context can improve recommendations for database queries. It uses deep learning to suggest useful queries while accounting for the patterns and needs represented in a user's existing workload.

Facilitating SQL Query Composition and Analysis (SIGMOD, 2020)

This work presents techniques for helping people construct and understand SQL queries. It focuses on supporting the full interaction around a query, from composing a statement to examining its behavior and results.

Improvement of SQL Recommendation on Scientific Database (SSDBM, 2019)

This paper investigates how SQL recommendations can be improved for scientific databases. The work considers the specialized workloads and query patterns found in scientific data analysis to make recommendations more useful to researchers.

A Generic Top-N Recommendation Framework for Balancing Accuracy, Novelty, and Coverage (ICDE, 2018)

This paper introduces a general framework for designing Top-N recommenders that balance several competing goals. In addition to accuracy, it considers whether recommendations are novel and whether they provide broad coverage of the available items.

Bridging the Gap Between User-Centric and Offline Evaluation of Recommendation Systems (UMAP, 2018)

This work examines the relationship between offline recommender-system metrics and the experience of real users. It highlights why evaluation should connect algorithmic measurements with user-centered outcomes rather than relying on a single offline score.

Facilitating User Interaction With Data (PhD@VLDB, 2017)

This research explores ways to make data systems more approachable for people who need to ask questions of data. It brings together query composition, recommendation, and answer analysis to support users throughout an interactive data exploration process.

Extracting Aggregate Answer Statistics for Integration (EDBT, 2015)

This paper studies how aggregate statistics about answers can be extracted and used to support data integration. The approach helps systems reason about query answers and combine information in a way that is useful for subsequent analysis.

Finding Expert Users in Community Question Answering (WWW Workshop, 2012)

This work investigates how to identify expert contributors in community question-answering archives. It uses signals from user activity and answered questions to help distinguish knowledgeable participants who can provide reliable guidance.

Modeling Community Question Answering Archives (NeurIPS Workshop, 2011)

This paper models the content and activity found in community question-answering archives. The goal is to better understand how questions, answers, and contributors relate to one another in these collaborative knowledge spaces.

Datasets

StackOverflow Q&A Archive Dataset

A dataset of 4,184 Stack Overflow questions and 15,822 question-answer pairs from the January 2011 data dump, plus 822 duplicate questions for testing. 21 tags were hand-selected using tag frequency and co-occurrence statistics, keeping the sample realistic while mixing overlapping and distinct topics; each question keeps up to its 4 highest-scored answers. For answer retrieval, the test set uses Stack Overflow's own community-identified duplicate questions (852 duplicate relations) as a gold standard, rather than requiring a user study. Built for the "Modeling Community Question Answering Archives" paper and used in my MSc thesis.

SDSS Query Workload Dataset

A SQL workload dataset extracted from Sloan Digital Sky Survey (SDSS) query logs: from 194 million log entries across roughly 1.6 million sessions, one query log was randomly sampled per session and identical statements were merged (aggregating their labels), yielding 618,053 unique query statements labeled with error class, CPU time, answer size, and session class. Built for the "Facilitating SQL Query Composition and Analysis" paper.

Posters and Demos

Facilitating Data User Interaction With Data (Microsoft Research AI Breakthroughs Workshop, 2019)

This demonstration presents an interactive approach to helping people work with data. It emphasizes practical user interaction and shows how research ideas can support more direct and productive data exploration.

Facilitating Data Exploration, Query Composition, and Query Answer Analysis (NWDS, 2018)

This demo brings together several stages of a data workflow in one user-facing experience. It supports exploring data, composing queries, and examining answers so that users can move more easily from an initial question to an informed result.

Personalized Top-N Recommendation for Promoting Long-Items (WIML/NeurIPS, 2017)

This work considers how personalized Top-N recommendation can help surface longer or less frequently selected items. It focuses on recommendation strategies that account for user interests while improving visibility beyond the most obvious short-list choices.

Talks


© 2026 Zeinab Zolaktaf