Zeinab Zolaktaf

Zeinab Zolaktaf

Vancouver, Canada

I'm an applied scientist building generative AI systems, from LLM evaluation and experimentation tooling to production applications, often across teams. Lately I've been using agentic AI to move much faster than I ever could on my own.


Experience

Amazon - Applied Scientist (2022–Present), Vancouver, Canada

Elpha Secure - Lead Applied Scientist (2021–2022), Vancouver, Canada

Olyns - Machine Learning Consultant (2021), San Francisco, USA

Georgian - Applied Data Scientist (2020–2021), Toronto, Canada

EhsAI - NLP and Machine Learning Scientist (2019–2020), Vancouver, Canada


Education

PhD, University of British Columbia (2012–2019), Vancouver, Canada

MSc, Dalhousie University (2009–2012), Halifax, Canada

BSc in Computer Software Engineering, Isfahan University (2004–2008), Isfahan, Iran


Publications

Workload-Aware Query Recommendation Using Deep Learning (EDBT, 2023)

This paper studies how workload context can improve recommendations for database queries. It uses deep learning to suggest useful queries while accounting for the patterns and needs represented in a user's existing workload.

Facilitating SQL Query Composition and Analysis (SIGMOD, 2020)

This work presents techniques for helping people construct and understand SQL queries. It focuses on supporting the full interaction around a query, from composing a statement to examining its behavior and results.

Improvement of SQL Recommendation on Scientific Database (SSDBM, 2019)

This paper investigates how SQL recommendations can be improved for scientific databases. The work considers the specialized workloads and query patterns found in scientific data analysis to make recommendations more useful to researchers.

A Generic Top-N Recommendation Framework for Balancing Accuracy, Novelty, and Coverage (ICDE, 2018)

This paper introduces a general framework for designing Top-N recommenders that balance several competing goals. In addition to accuracy, it considers whether recommendations are novel and whether they provide broad coverage of the available items.

Bridging the Gap Between User-Centric and Offline Evaluation of Recommendation Systems (UMAP, 2018)

This work examines the relationship between offline recommender-system metrics and the experience of real users. It highlights why evaluation should connect algorithmic measurements with user-centered outcomes rather than relying on a single offline score.

Facilitating User Interaction With Data (PhD@VLDB, 2017)

This research explores ways to make data systems more approachable for people who need to ask questions of data. It brings together query composition, recommendation, and answer analysis to support users throughout an interactive data exploration process.

Extracting Aggregate Answer Statistics for Integration (EDBT, 2015)

This paper studies how aggregate statistics about answers can be extracted and used to support data integration. The approach helps systems reason about query answers and combine information in a way that is useful for subsequent analysis.

Finding Expert Users in Community Question Answering (WWW, 2012)

This work investigates how to identify expert contributors in community question-answering archives. It uses signals from user activity and answered questions to help distinguish knowledgeable participants who can provide reliable guidance.

Modeling Community Question Answering Archives (NeurIPS, 2011)

This paper models the content and activity found in community question-answering archives. The goal is to better understand how questions, answers, and contributors relate to one another in these collaborative knowledge spaces.

Datasets

StackOverflow Q&A Archive Dataset

A dataset of Stack Overflow questions, answers, and duplicate-question pairs. Tags were deliberately selected by frequency and co-occurrence to balance easy and hard cases, and the test set uses Stack Overflow's own community-identified duplicate questions as a gold standard for answer retrieval, rather than requiring a user study. Built for the "Modeling Community Question Answering Archives" and "Finding Expert Users in Community Question Answering" papers and used in my MSc thesis.

SDSS Query Workload Dataset

A SQL workload dataset extracted from Sloan Digital Sky Survey (SDSS) query logs: sampled one query per session and deduplicated down to 618,053 unique query statements from an original 194 million log entries across roughly 1.6 million sessions. Built for the "Facilitating SQL Query Composition and Analysis" paper.

Posters and Demos

Facilitating Data User Interaction With Data (Microsoft Research AI Breakthroughs Workshop, 2019)

This demonstration presents an interactive approach to helping people work with data. It emphasizes practical user interaction and shows how research ideas can support more direct and productive data exploration.

Facilitating Data Exploration, Query Composition, and Query Answer Analysis (NWDS, 2018)

This demo brings together several stages of a data workflow in one user-facing experience. It supports exploring data, composing queries, and examining answers so that users can move more easily from an initial question to an informed result.

Personalized Top-N Recommendation for Promoting Long-Items (WIML/NeurIPS, 2017)

This work considers how personalized Top-N recommendation can help surface longer or less frequently selected items. It focuses on recommendation strategies that account for user interests while improving visibility beyond the most obvious short-list choices.

Talks


© 2026 Zeinab Zolaktaf