LinkedIn's Prompt Engineering Tool for AI Use in Production
Product Design
Tooling
Role
Design Lead
Timeline
Apr '23 (4 weeks)
Partners
1 Designer, 1 PM, 1 Frontend Engineer, 1 Backend Engineer
When LinkedIn set out to produce AI-generated collaborative articles at scale, the bottleneck wasn't the AI, it was the humans evaluating the output to see if it was good enough for our product. I designed the internal platform that changed that: a human review tool that helped ship over a million articles and became the foundational for how LinkedIn builds features with AI.
DISCOVERY
A New Frontier with AI Prompts
This project kicked off in 2023, early in the generative AI wave. Every team was scrambling to ship AI features and LLMs were new to most people, including ours. Before designing anything I researched the space independently by reading about LLMs and experimenting with model parameters like temperature and output length to build enough foundation to make good decisions.

The Process Behind the Articles
These articles needed AI-generated content published under LinkedIn's brand. The team making it happen was small, two prompt writers and five reviewers ensuring every article was accurate and on-brand I worked directly with them, sitting in on their process and creating journey maps to document and share their workflow with the broader team.
What we found through these conversations was a scrappy process in a slew of Google Docs, Sheets, and Slack Threads that worked but couldn't scale:
No version control
Prompt changes were tracked manually in a Google Doc.
No structured feedback Scores and notes were logged inconsistently and hard to consolidate.
Not scalable
Everything depended on manual coordination between a small team.
OPPORTUNITY
How might we give our prompt team a way to build and iterate on AI content at speed without losing control over quality and accuracy before it reaches members?
DESIGN
Defining Data Architecture with Eng
I worked closely with the PM and engineering team to define the data structure, wireframing in real time throughout. We ran continuous feedback sessions with the prompt writing and review team to make sure the structure matched how they actually worked.
The key decision was introducing a Project layer above individual prompts. For Collaborative Articles, there were multiple prompts working together each with its own review cycle but belonging to the same initiative. This led to a three-tier hierarchy with versioning built in.

Iterating in Real Time
Some of the more technically complex interactions were worked out directly on the whiteboard with engineering before any design work began. We ran continuous feedback sessions with the prompt writing and review team alongside this, making sure what we were building matched how they actually worked.
Two interactions required especially close collaboration: the dynamic variable system, which introduced a non-standard text editor pattern where writers could insert hooks directly into prompt text and upload a CSV to populate them at scale, and the LLM parameter controls, which needed to surface technical concepts like temperature in a way that was actionable for non-technical writer.
By the end of the sprint we had a clear enough direction (and a messy whiteboard) to move straight into high fidelity.
FINAL
Final Solution
Our final solution was an end-to-end platform that allowed teams to iterate on prompts quickly, with built-in feedback loops and a clear path to production.
Project and Prompt List Hiearchy
Each project contained multiple prompts with a full version history. Status tags showed whether each version was in production, in testing, or deprecated, giving teams a clear view of the prompt lifecycle at a glance.
Variable Configuration
When adding prompts, writers inserted hooks directly into the prompt text and uploaded a CSV to populate them at scale. A validation system confirmed each hook was matched to the correct CSV column before testing began. This was one of the trickier interactions to design, but the inline hook chips and real-time validation made a technically complex system feel intuitive for our writers.

Human Review Interface
Writers configured LLM settings like temperature and top-p before running test cases. Plain language descriptions and a warning state for high temperature values helped writers make informed decisions without needing technical expertise.

Summary Dashboard
An aggregated score gave teams a clear signal on whether to ship or iterate. Reviewer ratings fed back into model training over time, creating a feedback loop that improved output quality.

RESULTS
Outcomes
Genflow launched internally and was quickly adopted across LinkedIn. The editorial team used it to ship AI content at a scale that wouldn't have been possible with their previous spreadsheet process, and the platform went on to support other AI initiatives across the company.
1M+ Collaborative Articles Launched
Produced by the editorial team within a month of launch. Read more about Collaborative Article here.
Company Wide Adoption
The platform became the foundation for other AI tools across LinkedIn, including LinkedinNews, creator tools, profile updates, and more.
This project reinforced that the most impactful design work isn't always the most visible, sometimes it's the infrastructure that makes everything else possible.
READ MORE CASE STUDIES
Let's stay connected —




