vault backup: 2024-08-08 15:47:33

This commit is contained in:
2024-08-08 15:47:33 -05:00
parent 59d0938651
commit dc590594d2
492 changed files with 104 additions and 72 deletions

File diff suppressed because one or more lines are too long

View File

@@ -0,0 +1,45 @@
# Chapter One
The two parts of radical candor are: care personally, challenge directly.
Caring personally can be accomplished by "bringing your whole self to work". Bring your whole self to work means caring about people. It may mean being hated--even though you care deeply, personally about everyone on the team!
The hardest part about challenging directly is not giving it to others, but creating the environment where they can give it to you. This may make _you_ angry at _them_ and can be hard to deal with!
Being radically candid is not accomplished by simply starting your feedback with "let me be radically candid with you"! You have to care personally. It also does not mean saying everything you want to. A good rule of thumb is to leave three things unsaid every day.
Radical candor is not about endless communication that exhausts introverts. It's not about endless activities together to "build team". It's simply being human.
You may need to adjust what this means by country, team, company, etc. Not every "radically candid" practice will translate directly as you apply it in different contexts.
- In Israel, respect and caring personally was shown by being extremely direct and even aggressive in communicating with each other
- In Japan, it was more about being persistently polite
# Chapter Two
Creating a culture of open communication
Caring personally and challenging directly = radical candor
Challenging directly without caring personally = obnoxious aggression
Caring personally without challenging directly = ruinous empathy
Neither = manipulative insincerity
A useful quip: "it's not mean, it's clear". Being clear can _avoid_ being mean later--like firing someone after not being clear about their work early enough!
When praising, giving specific praise is best.
Most people for a "competent asshole" than a "nice incompetent". You don't have to choose between the two. You can be a competent nice manager.
Ruinous empathy doesn't work because it's aimed at making someone feel better instead of helping someone improve at their job.
# Chapter Three
You have to understand what motivates each person on your team--how does their job fit into their life goals?
In WWII the US Air force started to bring their best pilots home to train the next ones. This lead to finally defeating the Nazi air force because they tended to fly their aces until they were shot down.
I would say that Justin Strandburg is a "rock star" (rock of Gibralter type rock, not Guns 'n' Roses rock star) and that Ezekiel Pierson is more of the super star. Ezekiel is ambitious, rising quickly, learning a lot.
Rock stars need to be recognized as valuable to the company--they get a lot done--and not disrespected for being "b players". I should be careful to no try to promote them, if they are happy in their current jobs. We don't want to fall into the Peter principle where workers are promoted to the position of least competence.
Super stars need to be continually challenged.

View File

@@ -0,0 +1,102 @@
# Introduction
Xvii.
- Is this true from our observation of culture?
- Are we specifically in this boat too?
- What direction have we received?
- Are we applying it in our parenting today?
- Claims about our culture:
- Lost its way regarding parenting
- No sense of direction
- Can't direct ourselves
- Have children, don't want to be parents
- Need to quench personal thirst for fulfillment
- Children viewed as a liability
- Parents minimize time with children
- Quality over quantity time
Xviii.
- Is this our generation or our parents'?
- Written in 1995
- Protests in 60s led to rebellion against authority
- This led to cultural rejection of authority
- This affected parenting by destroying the "do as I say or else" method of parenting
- "the old ways of parenting no longer work" - did they ever? Was it effective or did it just lead to the rebellion of the 60s?
- Parents have responded in two ways
- Give up (parenting is impossible)
- Keep trying old "john wayne" methods
Xix.
- "the only safe guide is the Bible" - refreshing for me, when all other authoritative sources seem untrustworthy
- My thought: success in parenting, as with other things in life, is dependent on our view of God and his revelation (the Bible). If we believe it's His word and we can trust it, we have a solid way forward. If we don't, we have no way forward. We become the judges of how to parent. We first have to be humble to submit to His authority on the matter.
- "God's ways are no proved inadequete; they are simply untried". I feel like my parents, because they were first-generation, serious Christians, did go to Scripture and the Lord for biblical parenting.
- Was this everyone's experience?
- Or is what they saw more the "God's ways were untried" experience
- Biblical parenting
- Kind authority
- Shepherding children to understand themselves in God's world
- Keep gospel clearly in view
- Goal of someday living in mutuality
- Authority - nothing new, we all live under authority, whether parents or not. No surprise that parenting includes this aspect as well
- Are we embarassed to be authorities for our kids? If so, what in our culture makes it that way?
- Movie representation of parents?
- "scientific" methods of parenting?
Xx.
- "empower them to be self-controlled" - is this why authority is "embarassing"? It's not understood as an empowering tool?
- Jesus is our example in "kind authority" -- though God, He came to serve
- "children generally do no resist authority…" -- powerful statement. But important that it's a mix of truly kind and selfless authority
Xxi.
- "Values and spiritual vitality are not simply taught, but caught" - we need to be examples in deed, not just word
- how can we do this when we're tired, exhausted parents?
- Refresh ourselves
- Dates
- Church
- Friends
- Vacation
- Watch out for each other
- This book study
- Get together
- Mom's get-togethers
- Guys nights together
- Stay connected with the Lord
- Church attendance
- Quiet time
- Sacrifice! Its not easy
Xxii
- Not simply well-behaved (external)--the gospel works from the inside out
- Gospel enables accurate picture of yourself and God
- I'm a sinner who fails
- God is righteous who saves with great grace
- Keepable standard
- Because they're not christians?
- I do it because it's easier! Not good….
- Impossible standard forces facing the gospel--powerful statement, never thought of it this way
Xxiii.
- "Each child will examine" - how else can we prepare? It will happen, give them the tools to do it well
- Mutuality as people under God - seems a long way off, but Im looking forward to it, what a reward

View File

@@ -0,0 +1,94 @@
Clipped from: [https://www.parand.com/a-completely-non-technical-explanation-of-ai.html](https://www.parand.com/a-completely-non-technical-explanation-of-ai.html)
# Overview
This document will explain what neural networks are and how they work, which will help you understand how AI and machine learning work. In the scenario below you'll play the part of the neural network.
# Day One
First day of your new job as a "classifier" your boss walks in and drops a big spreadsheet of numbers on your desk.
"Is this a cat?" she asks.
Confused you ask "What?"
![Exported image](Exported%20image%2020240808113913-0.png)
"Is this a cat?"
Even more confused, you respond "I don't know".
"Wrong!" she says, and slaps you across the face.
Before you've had a chance to be shocked she drops another large spreadsheet on your desk. "Is this a cat?"
"I don't understand" you respond.
"Wrong!" she says, and slaps you again.
Another spreadsheet. "Is this a cat?"
Not wanting another slap, you meekly respond "Yes?"
"Correct! Good job!" she says, and gives you a wonderful reward. Almost makes up for the slaps.
Another spreadsheet. Cat?
Slightly more confident and wanting another reward you respond "Yes"
"Wrong!". Another slap.
You are very confused. "I don't understand what's going on. You haven't told me the rules, you haven't told me how to figure out something is a cat, you haven't given me the logic to figure this out. You haven't trained me."
"Correct" says the boss. "This is training. Trust me, you're going to get very good at recognizing cats"
![Exported image](Exported%20image%2020240808113913-1.png)
Another spreadsheet. This time you focus on the sheet. It's 256 columns and 256 rows, filled with numbers. Is it a cat? You don't know, so you guess.
Another and another. Right, wrong, wrong, right again, it keeps going. Slaps and rewards.
You're starting to notice some patterns - if the sheet is almost all zeros then it's not a cat. You look at the center of the sheet - that part should have larger numbers.
It's getting slightly better - you're getting more rewards than slaps, guessing correctly more often than not. It's all about the patterns of the numbers.
# Day Two
This is the strangest job you've ever had. The slaps are terrible, but the rewards are so great you don't want to quit. How do you get better at this? It's not really possible for one person. If you had more people they could focus on different aspects of the sheet, look for different patterns, and you could use their findings to make better guesses.
You hire 10 people and bring them to work the next day. When you get a spreadsheet you show it to those 10 people and ask them "Is this a cat?"
They are as confused as you were. You force them to guess. The first guy looks like an idiot, so you decide to go with the opposite of what he says. The third lady looks really thoughtful so you put a lot of weight on what she says. In your mind you assign a weight to each of their guesses to come up with your final answer.
Each time you get a reward or punishment you share it with the 10 people you've hired: you reward or slap them based on how much weight you put on their input and how much they contributed to you getting the answer right or wrong. You learn the patterns: if the first guy says a strong no and the third lady a strong yes then it's very likely a cat. You learn many more patterns like this.
The 10 people are learning to hone their opinions based on the rewards and punishments they get. They can focus on different aspects of the sheet, and their opinions together can tell you a lot about the sheet.
After what feels like an epoch and a lot of spreadsheets, eventually you get better. You're getting more rewards. This is working. Your boss says you're very perceptive, and starts calling you Perceptron.
![Exported image](Exported%20image%2020240808113913-2.png)
# Day Ten
What if you get more people involved? You could make it 50 people reporting to you, but that'd be hard to manage. How about we add another layer of people before your 10?
You hire 50 more people, have them give their guesses to the 10 people that report to you. You're now very removed from looking at the spreadsheets - instead you rely on the patterns found by the first layer of the people, who give their opinions to the second layer of the people, who then inform you. Everybody passes the slaps and rewards down the line based on how much each earlier person's guess contributed to their guess.
It takes even longer, but eventually the system starts working. You're much more accurate in identifying when it's a cat.
Interestingly you were never taught the rules or logic, and you didn't teach your people the logic or rules. You just propagated the rewards and punishments back through each layer: the more each person's opinion contributes to the answer, the more rewards or punishment you shared with them, and they in turn with the people in the layer behind them. Each person uses the same method to propogate their rewards and punishment back to the people in the layer below them.
# Neural Networks
This is how neural networks work: they see many examples and get rewarded or punished based on whether their guesses are correct. They use multiple layers of workers and eventually learn patterns. Importantly no one is teaching them what patterns they should be looking for or telling them the logic or the rules - the networks eventually figure out the patterns and logic based on very many rounds of example, reward, and punishment. This is called **machine learning**, because the machine is learning the rules by itself.
Neural networks work well when you have many examples of something (eg. pictures of cats), but it's hard to write the logic and rules to describe how to recognize that thing. Try it - write down some rules for how to recognize a cat (eg. "has 4 legs"), then look at pictures of cats and see where the rules fail (eg. a picture of a cat's head). In these cases you use machine learning so the machine learns the rules by itself.
Also note that computers see things as multi-dimensional tables of data. They don't look at a "picture" - they see 3 spreadsheets of numbers representing the RGB values of the picture.
# Day Forty
Identifying cats is a lucrative business and you're pretty good at it. But not good enough. How can we get even better? More people, more layers!
Unfortunately it takes a long time for people to do their calculations, so you've been stuck with just a few layers of people for a while. Having a lot more people would make the process take too long.
One day you run into a group of people who call themselves Gaming People United (GPU). These people play a game that has taught them to be very good at looking at spreadsheets and calculating numbers - in fact they can do a lot of calculations in parallel, very quickly.
Excited, you hire a bunch of them and put them to work recognizing cats. Now instead of two layers of people you can have 10! Each layer can focus on higher and higher concepts - the first layer can look for small details (eg. do I see a pattern that looks like a small circle? Do I see a pattern that looks like a sharp edge?), the second layer can look for patterns from the results of the first layer (eg. are there two circles close to each other), the third layer can build on that (eg. are there two circle close to each other, and a triangle below them), and so forth. By the time the results of the layers get to you you have some fairly sophisticated concepts - we see a pattern that looks like a face, we also see patterns that could be legs, and a sharp thing that might be a beak. Now your guess is more informed than ever.
You train the layers sending slaps and rewards down through each layer, and after many epochs you get really good at recognizing cats.
![Exported image](Exported%20image%2020240808113913-3.png)
# Deep Learning
One of the major break-throughs in machine learning was the advent of _Deep Learning_, which is basically what we describe above - GPUs (Graphics Processing Units) got popular because they enable fast 3D graphics for games. They also happen to be very good at quickly doing the types of calculations neural networks need. Machine learning people started using GPUs for training neural networks, and with this extra speed they could have many more layers - from 3 or 4 layers to 10 or 11, and then hundreds.
This addition of layers led to significant advances in neural network performance - very quickly neural networks became the best solution for voice recognition, for image recognition, image creation, and sophisticated language models. This is called **Deep Learning** because there are so many layers, not because it's profound.
_You can stop reading here if you want, the following portion is only here because someone asked me what "convolutional neural networks" are._
# Day Fifty
How can we make the process even more efficient? You start to realize that looking at the entire sheet is too hard - the first layer of people literally have to stare at the whole thing and try to guess based on that. What if we give them a small portion of the sheet to look at, give a guess for that portion, then move on to the next portion of the sheet, and so on? The first layer of people are focused on lower level concepts anyway - looking for edges, things that look like circles, and so forth - and they can find those in the small portions of the sheet, without the need to look at the whole thing at once.
![Exported image](Exported%20image%2020240808113913-4.gif)
[Via towardsdatascience](https://towardsdatascience.com/intuitively-understanding-convolutions-for-deep-learning-1f6f42faee1)
You tell the first layer of people this is what they should start doing. They respond that this sounds very convoluted and they'll only agree to do it if you call them "colonel".
This makes the process even faster and more accurate: the first layer become specialists in small features of the sheet, and they learn to be very efficient and fast. You try the same idea for the next several layers: you tell people to focus on small portions of the feedback they get from the layer below them, moving the portion they focus on so eventually the scan across all of the feedback.
This specialization and focus makes you even more accurate. Congratulations, you are officially recognized as the best cat classifier in the world.
# Convolutional Neural Networks
Convolutions are this idea of applying a specific set of calculations, sometimes called a kernel, to each portion of the input, and scanning your window of attention across the entire image. The set of calculations, or kernel, is learned by each worker - you don't tell it what to calculate or how, it learns what's useful based on rewards and punishments. Convolutional Neural Networks (CNNs) were a significant step forward in the capability of neural networks.
# Further Reading
_If you're interested in this topic you might also enjoy the_ _Completley Non-Technical Explanation of ChatGPT_ _series as well._

View File

@@ -0,0 +1,77 @@
Clipped from: [https://www.parand.com/a-non-technical-explanation-of-chatgpt.html](https://www.parand.com/a-non-technical-explanation-of-chatgpt.html)
Continuing the series of non-technical explainers, let's figure out how ChatGPT (and in general Large Language Models, or LLMs) work. This is part one of the ChatGPT explainer, with part two coming soon.
It's helpful but not necessary to read the [non-technical explanation of AI](https://www.parand.com/a-completely-non-technical-explanation-of-ai.html) first, if you feel like it read that and come back.
# Fill In The Blank
Let's play a quick game of fill in the blank:
to be or not to _____
Why did you immediately think of the word _be_ to fill in that blank, instead of the word _banana_ or _fish_? Because you've seen the phrase many times and your brain has learned the most likely next word is _be_.
Let's do another one:
rock and ____
Did you think of the word _roll_? Why? Because that's the word you see most often following the words _rock and_.
How about this one:
I am very _____
This is less clear - the next word depends on the context.
I just ran a marathon. I am very _____
Perhaps now you would say _tired_. Or _happy_, or _proud_.
Your brain is picking the most likely next word based on the context of the sentence and based on what words it's seen most frequently in that context.
# Aliens and CatGPT
In an astonishing turn of events a race of alien cats have landed on earth. Since you are [the world's foremost cat expert](https://www.parand.com/a-completely-non-technical-explanation-of-ai.html) you have been chosen to communicate with them.
![Exported image](Exported%20image%2020240808113915-0.png)
Your boss walks in and drops a massive print out of all the alien cat communications on your desk.
"Speak cat!" she commands.
Unfortunately you don't speak alien cat.
You page through the print outs - it looks like pages and pages of gibberish.
zoog zeeg zag. zoog zeeg kaz. zoog zeeg bah. rag zoog zeeg. rag zoog ko. kaz rag. zap zoog zeeg. ...
# Markov to the Rescue
What to do? You remember your friend Markov who's always talking about languages, words, and their relationships. Maybe he can help. You show him the alien cat communications and ask him if he can help you say something in cat.
"Yes!" he exclaims - "We can do this. We will create our response one word at a time, just by picking the best word to say next, and keep going until we have a sentence"
"But how do we know what the best word to say next is? Don't we need to know what the words mean?"
"Nope, we just need to know what word to say next"
"That... doesn't seem like it would work"
"Let me show you", says Markov, and grabs the book on his desk, [Alice in Wonderland](https://www.gutenberg.org/ebooks/11).
![Exported image](Exported%20image%2020240808113915-1.png)
"We just need to know what word to say next. The best word to say next is simply the word that shows up most frequently after the word we're looking at. All we have to do is make a table of how often each word follows another word. We'll call this a frequency table"
He shows you how to create the frequency table - you just note down how many times each word follows another. It takes a while to go through all Alice in Wonderland and do the counting. You end up with:
Next word after "Alice" → was: 17 times, and: 16 times, thought: 12 times, had: 11 times, ...Next word after "sat" → down: 9 times, silent: 2 times, still: 2 times, for: 1 times, ...Next word after "was" → a: 31 times, the: 20 times, not: 12 times, going: 11 times, ...Next word after "and" → the: 80 times, she: 52 times, then: 31 times, was: 19 times, ...Next word after "a" → little: 59 times, very: 25 times, large: 20 times, great: 17 timesNext word after "very" → much: 10 times, soon: 7 times, curious: 6 times, glad: 5 times...
This table gives you a lot of help in forming sentences in the style of Alice in Wonderland. For example, if you start with the word _Alice_, then you'd look that up in the above table and see that the most frequent next word is _was_. And after _was_ you would select _a_. After _a_ you would get _little_. You'd end up with _Alice was a little_.
How well does this work? Let's create some sentences:
the door, when i beg your verdict," it was quite plainly through the bottle, i'm afraid that they had its nest.
You can try this out for yourself on [almeopedia](http://almeopedia.com/markovtest.html).
That's not great. Why is it so nonsensical? Think back to our fill in the blank examples at the start of this post: in order to pick a good next word you need context. If someone asked you what word should appear after _rock_, you'd have a hard time picking something reasonable, but if they gave you _rock and_, you'd fairly quickly think of _roll_.
What would happen if you considered two words instead of a single word for your context? For one thing your frequency table creation would become harder - now instead of a single line for each word, you'd have a line for each word pair. You'd have to gather stats for all the two word permutations, which is a lot more than the single word case.
Next word after "Alice was" → not: 3 times, beginning: 2 times, very: 2 times, ...Next word after "was not" → a: 3 times, here: 1 times, going: 1 times ...Next word after "not here" → before: 1 timesNext word after "not a" → moment: 2 times, bit: 2 times, serpent: 2 times ...
Let's create a sentence with this and see how it looks:
the hatter, and, burning with curiosity, she decided on going into the garden.
You can try this out for yourself on [almeopedia](http://www.almeopedia.com/markov3test.html).
That's looking better, it almost makes sense. Let's keep going - instead of two words of context, how about three?
The Hatters remark seemed to have no sort of chance of her ever getting out of the water, and seemed to quiver all over with diamonds, and walked two and two, as the soldiers did.
Four?
So she swallowed one of the cakes, and was delighted to find that she knew the name of nearly everything there.
Hmm. This is a little too good. It turns out a very similar sentence exists in the original Alice in Wonderland text:
So she swallowed one of the cakes, and was delighted to find that she began shrinking directly.
Our simple method of picking the most likely next word can result in the system memorizing the text snippets - given long enough context, the next most likely word is exactly the word that appeared in the original text following that context.
We can fix this by introducing some randomness - instead of always picking the most likely next word, we can pick somewhat likely next words. This way we're less likely to regurgitate the original text.
The good news is our method seems to work for English, so it'll probably work for alien cat language as well.
_Aside: You might run into the term "stochastic" as you look into language models - this just means randomly determined. For example, you might hear people argue whether these systems are_ **Stochastic Parrots**_, implying the systems are simply parroting back the original text that they saw, with some randomness thrown in._
# Large Language Models
ChatGPT and many other Large Language Models (LLMs) do essentially what you did above: they examine a very large amount of human communication, gather stats (or probabilities) on what words are most likely to follow other words, and play a continous game of fill-in-the-blank. You give them some context (your prompt or question), and they create the reply one word at a time by selecting the word most likely to appear next. They respond with a word, then look at their internal stats to pick a word to follow their first word, then a word to follow their second word, and so forth, one word at a time, until they've formed a response.
# To be Continued...
In [part two of this series](https://www.parand.com/a-non-technical-explanation-of-chatgpt-deep-learning.html) we'll look at the problems you'll run into with this method and how deep learning helps you overcome those problems.
# Further Reading
_If you're interested in this topic you might also enjoy the other posts in the_ _Non-Technical Explainer_ _series as well._

View File

@@ -0,0 +1,193 @@
Clipped from: [https://ferd.ca/embrace-complexity-tighten-your-feedback-loops.html](https://ferd.ca/embrace-complexity-tighten-your-feedback-loops.html)
This post contains a transcript of the talk I wrote for and gave at [QCon New York 2023](https://qconnewyork.com/presentation/jun2023/embrace-complexity-tighten-your-feedback-loops) for [Vanessa Huerta Granda](https://qconnewyork.com/speakers/vanessahuertagranda)'s [track on resilience engineering](https://qconnewyork.com/track/jun2023/resilience-engineering-culture-system-requirement).
The official talk title was "Embrace Complexity; Tighten Your Feedback Loops". Thats the descriptive title for the talk that follows the conferences guidelines about good descriptive titles. Instead I decided to follow my gut feeling and go with what I think really explains my perspective and the approach I bring with me to work and even my life in general:
![Tkts Is ALL GOING To KELL ACL WE CAW DO IS (ÄFL(JENC€ Pow CT's TAKE frea Hebert keeps ](Exported%20image%2020240808113916-0.png)
I take what would probably be a sardonic approach to dealing with life and systems, and so “This is all going to hell anyway” is pervasive to my approach. Things are going to be challenging. There are going to always be pressures that keep pushing our systems to the edge of chaos. I dont think this can be fixed or avoided. Any improvement will be used to bring it right to that edge. In complex systems, the richness and variability is often there for a reason. Trying to stamp it out in favour of stronger control is likely to create weird issues.
So the best I personally hope for is to have some limited influence in steering things the best I can to delay going to hell as long as possible, but thats it. And my talk is going to focus on a lot of these approaches, but first, I want to explain why I feel things are that way.
![OFF TE Figure 2. (Color online) Process Map From: - OF -tl-e ](Exported%20image%2020240808113916-1.png)
In what is probably my favorite paper ever, titled [Moving Off The Map](https://ferd.ca/notes/paper-moving-off-the-map.html), Ruthanne Huising ran ethnological studies by embedding herself into projects within many large corporations doing planned organizational changes. In supporting these efforts, they were doing “tracing” of their functions, which meant gathering a lot of data about what activities take place, what interactions and hand-offs exist, what information and tools are used and required? How long do tasks take? How do people and teams deal with errors? Generally asking the question “what do we do here?” and wondering with whom they do it.
To build these maps they generally reached out to experts within the organization who were supposed to know how things were working. Even then, they were really surprised.
![Figure 2. (Color online) Process Map OFF TE k(it was (ike the sun rose for the first time... r sau the bigger picture." k(Tke proble%f is that it was not designed in the first place." k(Tkis is even Biore fucked up than r imagined." ](Exported%20image%2020240808113916-2.png)
One explained that “it was like the sun rose for the first time… I saw the bigger picture.” Participants had never seen the pieces (jobs, technologies, tools, and routines) connected in one place, and they realized that their prior view was narrow and fractured, despite being considered experts.
Others would state that “the problem is that it was not designed in the first place.” The system was not designed nor coordinated, but generally showed the result of various parts of the organization making their own decisions, solving local problems, and adapting in a decentralized manner.
The last quote comes from events when a manager at one of the organizations walked the CEO through the map, highlighting the lack of design and the disconnect between strategy and operations. The CEO sat down, put his head on the table, and said, “This is even more fucked up than I imagined.” He realized that the operation of his organization was out of his control, and that his grasp on it was imaginary.
![CONTROL Pgoc.ESs COSTS SAVINGS I LLUSoRY To coczECT€b DATA CREhTED nap ](Exported%20image%2020240808113916-3.png)
One of the most surprising results reported in there was about tracking the people who participated in organizing and running the change projects, and seeing who got promoted, who left, and who moved around the org or industry they were in.
She found out there were two main types of outcome. The first group turned out to be filled with people who got promotions. They were mostly folks who worked in communications, training, who managed the costs and savings of the projects, or those who helped do process design. Follow-up interviews revealed that most of them attributed their promotions to having worked on a big project to put under their belt, and to frequently working with higher-ups, which both helped with getting promoted.
Another group however mostly contained people who moved to the periphery: away from core roles at the organization, sometimes becoming consultants, or leaving altogether. Those who fit this category happened to be the people who collected the data and created the map. They attributed their moves to either feeling like they finally understood the organization better, felt more empowered to change things, or became so alienated by the results they wanted to get out.
So the question of course became how come people who feel they understand how the organization truly works and who want to change it move _away_ from the central roles and positions, and into the peripheral ones?
!["FATAL «organizations and institutions exist on(g actual people's doings and that these are necessari(g particular, local and epkæeral" «socia( worlds do not have independent, stable existence but instead emerge from our collective action" «t.Uitkin the system of roles, rules, and routines, there is far %tore room to htaneuver than previous(g assumed." ](Exported%20image%2020240808113916-4.png)
The fatal insight, according to Huising, is something sociologists knew for a good while: the culture and the order imposed to organizations, groups, and even societies is often emergent and negotiated. And while it's obvious that these structures dictate a lot of actions, the actions themselves can preserve or change the structures around them.
The feelings of empowerment and alienation come in no small part because people realized that they could change a lot more than they could, albeit often from outside the core decision-making that enforces the structure (while understanding how that core works), or because the ways they thought they were impacting things was shown not to be effective and they felt disembedding.
![NOetstNAL Vs. 0-0-0 0—0-0 o o o ](Exported%20image%2020240808113916-5.png)
Another thing you have possibly experienced and isnt in the paper now is one of differentiating between the nominal and actual structure of the org, the emergent one that depends on power dynamics, who knows what or whom, who likes or dislikes each other, and so on.
If you've ever worked in a flat organization, like the one in the middle here, is that even though you have little management structure to speak of, power dynamics and decision-making authority still exists. People who have no power attached to their role are still going to be consulted or inserted in the decision-making flow of the organization, they're still going to be influential and have the ability to make or break projects, but just with less obvious accountability.
The nominal structure is the one where each level of management and within the organizational ladder specifies how information flows, and how authority is applied. It's what we see on the left in a more traditional org structure, and this way of organizing groups will simultaneously be useful to align efforts and to constrain them. It makes accountability more explicit and transparent, but structurally will prevent people from doing unspecified things, whether they would be harmful or useful.
The emergent structure is always there as well. It is implicit, always changing, and not necessarily constrained to your own organization either. Sometimes, people who know how to run, maintain, or operate components, or whom people listen to, are not even in your org anymore. They might have moved away (to a different team or even a competitor), retired, or never been in and they have just published a really influential piece of media and people look up to them.
But who knows what, works with whom, and who can move things around in specific contexts can be key to successful initiatives. Even if the organizational structure has often been put in place to constrain change, as a barrier to people working in mis-aligned ways, some folks central to the emergent structure, in key contexts, have earned enough trust to be allowed tacitly to bend and break the rules. They can choose not to enforce the rules, or the rules are not enforced as tightly for them with the hopes of positive outcomes—even if sometimes it can get you the opposite result.
Im not here to argue in favor of one or the other structure, but mostly that in my experience, driving change or making initiatives succeeds the most when catering to both structures at once, or rather fails when only looking at one and being blocked by the other. They're both real, both distinct, and pretending only either exists is bound to cause you grief.
![THE GAP ETwEEN WORK-AS-... JRK- ms- IMAGINED unRk-AS-DscLSED WORK-AS-DONE ](Exported%20image%2020240808113916-6.png)
As a continuation of this, the way people work every day is often different from the way people around them imagine their work is being done. The gap between how work is thought to be done and how it is actually done is a major but generally invisible factor in how systems work out.
Based on flawed mental models of the work, procedures and prescriptions are given about how to do work, and will vary in inaccuracy. People will imagine things like, for example, writing all the tests before writing or modifying any code and that code coverage could be ideal and then that it will all be reviewed in depth by an expert, and will enshrine this as a policy.
But the application of these policies is never perfect. Sometimes code doesn't have an owner, or due to crunch time and based on how much the reviewer and author trust each other, the review won't be as in-depth as expected.
When you see this mismatch causing people to ignore or bend rules, you can choose to apply authority and ask for a stricter rule-following. This pattern of enforcing the rules harder will likely drive these adaptations underground rather than stamping them out, because real constraints drive that behavior.
In turn, the work as disclosed will be less adequate, and the work as imagined progressively gets worse and worse.
This becomes a feedback loop of misunderstanding and at some point, like our devastated CEO, youre not managing the real world anymore.
![Fr Hebert mononcqc@hachyderm.io If you're a software developer who ever worked for an employer who had you track time hourly into specific projects/customer accounts and you were short on time budget, did you: ESCRO 13K¯ work for free/untracked/stopped work¯ 13% enter time in unrelated projects with more buffer 35% enter time in the same project regardlesE 58% - my time tracking was always fake and lies Refresh • 120 people • Closed Jan 21, 2023, 14:05 web ](Exported%20image%2020240808113916-7.png)
To demonstrate this, earlier this year I went to my local mastodon network—so you know this is super scientific—and ran a poll about time sheets. The question was "If you're a software developer who ever worked for an employer who had you track your time hourly into specific projects/customer accounts and you were short on time budget, did you..."
Multiple answers were accepted. Fewer than 15% of people either stopped work, worked without tracking their time anymore (for free), or shifted their time into other projects with more buffer space.
Roughly a third of people reported billing anyway, some stating that it's not their problem the time allocation wasn't realistic or adequate.
But the vast majority of answers, nearly 60%, came from people saying "my time tracking was always fake and lies," with some people stating they even wrote applications to generate realistic-looking time sheets.
What we can see here is an example of how work-as-imagined gets translated into policies ("people do their work in projects, and account for their time"), which at some point doesn't get applied right anymore. If I were to suppose, it could be things like not being allowed to go over time, or just finding the practice useless. But the end result is that the time sheet data just isn't trustworthy, and then it can get used again and again in further decision making.
The gap widens, and our CEO might also get to think "this is all fucked up."
![PRESSURES AO CONFLICTS WoeKLoAD RISK TRuST SOCCESS FAILURE ](Exported%20image%2020240808113916-8.png)
Part of the reason for this is that every day decisions are made by trying to deal with all sorts of pressures coming from the workplace, which includes the values communicated both as spoken and as acted out. People generally want to do a good job and theyll try to balance these conflicting values and pressures as well as they can.
The outcome of that trade-off being a success or a failure isnt known ahead of time, but these small decisions accumulate based on the feedback we get from each of these and can end up compounding and accumulating, either as improvements, or as erosion that makes organizations more brittle, or really anywhere in between. People adopt the organizations constraints as their own, and this set of pressures is the kind of stuff that drives processes to the edge of chaos over and over again.
These accumulations of small decisions, these continuous negotiations, thats one way your culture can define itself. Small common everyday acts and small amounts of social pressure you can apply locally has an impact, as minor as it might be, and compounds. You can easily foster your own local counterculture within a team if you want to. This can both be good (say in Skunkworks where you bypass a structure to do important work) or bad (normalizing behaviors that are counterproductive and can create conflict).
![Mow To EMBRACE GMPLFXITY? QADE-OFFs FVk1kJG KEEP EEDBACk LOPS HAC— — H OVER ](Exported%20image%2020240808113916-9.png)
So while a lot of the work you can do to improve reliability or resilience as a whole can be driven locally, my experience is that you nevertheless get the best results by also aligning with or re-aligning some of the organizational pressures and values usually set from above.
The idea here is to start looking at the organization from both ends: how can we support the people dealing with the trade-offs in conflicting goals as they happen, how can we influence the higher-level values and pressures such that we can try to reduce how often these conflicts happen even though they will definitely keep happening, and how can we better carry context and feedback across both ends so that we constantly adjust as best as we can. A system perspective on interactions, rather than focusing on components is also something I've found useful. The rest of the talk is going to be spent on these ideas.
_(as a note, the third drawing is_ _Dimethylmercury__, a highly volatile, reactive, flammable, and colorless liquid. It's one of the strongest known neurotoxins, and less than 0.1 mL is enough to kill you through your skin, and gloves apparently do a bad job at protecting you)_
![Mow To EMBRACE GMPLFXITY? QADE-OFFs FVk1kJG KEEP EEDBACk LOPS HAC— — H OVER ](Exported%20image%2020240808113916-10.png)
So let's start with negotiating trade-offs, with a bit more of an ops-y perspective, because that's where I'm coming from.
![Dou'T DELWEÄ WHAT Asm To ](Exported%20image%2020240808113916-11.png)
This is a painful one sometimes, especially when you have highly professional people who take their jobs seriously.
Locally for you as a DevOps or SRE team, there is a need for the awareness of what the organization and customers actually care about. Some availability targets become useless metrics because theyre disconnected from what users want, and youre just going to burn people out doing it.
I learned this lesson when talking to the SRE manager of one of these websites where people pick their favorite images, put them on boards, and get shown ads. He was telling me how their site was having a lot of reliability issues. It would keep going down, his team would do heroics to bring it back up, and it'd open all over again.
He felt his team was burning out. They were losing people, and their call rotation was so painful they were also having issues hiring back into it. He was seeing the death spiral happening and was wondering what to do.
He added that there were perverse incentives at play: every time the site went down, they stopped showing images, but not ads. That meant that during incidents, they still earned money, but no longer paid for bandwidth. The site was more profitable when it failed than when it worked, and seemingly, users didn't mind much.
They were not getting help, nobody seemed to consider it a problem. Not really knowing what to say, I just asked off-hand: "are you trying to deliver more reliability than people are asking for? What if you just stopped and let it burn more and rested your people?" He thought about it seriously, and said "yeah, maybe."
I never actually found out what happened after this, but it still stuck with me as a really good question to ask from time to time.
![DoucT DELWEP- WHAT To ](Exported%20image%2020240808113916-12.png)
In some cases, the answer will be "yes, we want to be this reliable". But you just won't be given the right tools to do it.
At Honeycomb, we want on-call rotations to have 5-8 people on them because thats what we think gives a good pace that maintains a balance between how rested and how out-of-practice people can be. Not too often nor not often enough.
But many services are owned by smaller teams of 3-4 people. If we wanted rotations to be made of people who know all their components in depth, where they could build expertise and operate what they wrote, we couldn't reach a sustainable frequency.
Instead, to keep the pace right, we tend to put together rotations made of multiple teams, for which people wont understand many of the components they operate. This in turn makes us prepare to deal with more unknown: fewer runbooks, more high-level switches and manual circuit breakers to gracefully degrade parts of the system to keep it running off-hours, and with different patterns of escalation.
We started leaning more heavily on this when a big public product launch required shipping a new feature, which was to be operated by a team that didn't have full time to get it operationally ready. When our SRE team was discussing with them what still needed to be done, we asked for a few simple things: a way to switch the feature off for a single customer, and a way to turn it off entirely, that wouldn't break the rest of the product. The rest we could add as we went.
We ended up using these switches a few times, one of which prevented a surprising write-amplification bug that could have killed the whole system, and instead let us wait a few hours for the code owners to get up and fix it at a leisurely pace. We're going to accept a bit of well-scoped, partial unavailability—something that happens a lot in large distributed systems—in order to keep the system stable.
The person wearing the pager often does triage and that weird issues will eventually be handled by code owners, just not right now.
This approach means that rather than working impossible hours and making inhuman efforts foreseeing the unforeseeable, we keep moving rather fast, gather feedback, find issues, and turn around a bit more on a dime. In order to do this though, theres a general understanding that production issues may turn parts of the roadmap upside down, that escalations outside of the call rotation can disrupt project work, and so on.
Thats one of the complex trade-offs we can make between staffing, training/onboarding, capacity planning, iterative development, testing approaches, operations, roadmap, and feature delivery. And you know, for some parts of our infra we make different decisions because the consequences and mechanisms differ.
![THERE ARE NO SUbSTtT(ffES Foe SAFETY ](Exported%20image%2020240808113916-13.png)
To make these tricky decisions, you have to be able to bring up these constraints, these challenges, and have them be discussed openly without a repression that forces them underground.
One of my favorite examples is from a prior job, where one of my first mandates was to try and help with their reliability story. We went over 30 or so incident reports that had been written over the previous year, and a pattern that quickly came up was how many reports mentioned "lack of tests" (or lack of good tests) as causes, and had "adding tests" in action items.
By looking at the overall list, our initial diagnosis was that testing practices were challenging. We thought of improving the ergonomics around tests (making them faster) and to also provide training in better ways to test. But then we had another incident where the review reported tests as an issue, so I decided to jump in.
I reached out to the engineers in question and asked about what made them feel like they had enough tests. I said that we often write tests up until the point we feel they're not adding much anymore, and that I was wondering what they were looking at, what made them feel like they had reached the points where they had enough tests. They just told me directly that they knew they didn't have enough tests. In fact, they knew that the code was buggy. But they felt in general that it was safer to be on-time with a broken project than late with a working one. They were afraid that being late would put them in trouble and have someone yell at them for not doing a good job.
When I went up to upper management, they absolutely believed that engineers were empowered and should feel safe pressing a big red button that stopped feature work if they thought their code wasn't ready. The engineers on that team felt that while this is what they were being told, in practice they'd still get in trouble.
There's no amount of test training that would fix this sort of issue. The engineers knew they didn't have enough tests and they were making that tradeoff willingly.
![THERE so NON-TECHNICAL WAYs Fop THINGS To GO (T ts SOt0ETlES åkAY foe TO TRADE OF THAT RI Sk ](Exported%20image%2020240808113916-14.png)
_(note: this slide was cut from the presentation since I was short on time)_
Speaking of which, sometimes its also fine to drop reliability because there are bigger systemic threats.
Sometimes you can eat downtime or degraded service because its going to keep your workload manageable and people from burning out. or maybe you take a hit because a big customer that makes you hit your targets as an org and can prevent layoffs will put some things over the limit and a components performance will suffer. You cant be the department of “no” and that negotiation has to be done across departments.
Conversely however, you have to be able to call out when your teams are strained, when targets arent being met and customers are complaining about it. It means you might be right, and some deadlines or feature delivery could be deferred to make room for others.
How do you deal with capacity planning when making your biggest customer renew their contract prevents you from signing up another one thats as big? Very carefully, by talking it out by all the involved people.
And sometimes that trade-off is very reasonable. And good engineering requires you to move it earlier in the lifecycle of software than just around incidents. Its much simpler to change the shape of a products features than it is to deliver the perfect distributed systems sometimes. Making your features take the ideal shape to deal with the reality of physics is one of the things a good collaborative approach can facilitate.
![Mow To EMBRACE COMPLEXITY? QADE-OFFs FVk1kJG KEEP EEDBACk LOPS HAC— — H OVER ](Exported%20image%2020240808113916-15.png)
So we can make tradeoff negotiation simpler by having these honest discussions, but in many cases this ability to discuss constraints to influence how work takes place brings us to this next step, where we dont only influence the decisions people make, but surface these challenges to influence how the organization applies its pressures. This is moving from the local level to the alignment to the broader org structure.
![METRICS APE TREE To You, BJOT YW To SERVE THOA ](Exported%20image%2020240808113916-16.png)
Metrics are good to direct your attention and confirm hypotheses, but not as a target, and theyre unlikely to be good for insights. [Theyre compression, and it can be unreliable](https://ferd.ca/plato-s-dashboards.html).
The thing you generally care about is your customer or user's satisfaction, but there's a limit to how many times you can ask "would you recommend us to a friend?" and still get a good signal. So you start picking a surrogate variable.
You assume that when the site is down and slow, people are mad, and you make being up and fast a proxy for satisfaction. But then that signal is a bit messy and not super actionable, because it can include user devices or bits of the network you don't control, plus it's hard to measure, so you'll settle for response time at the edge of your infrastructure. This loses fidelity into the signal, but it'll get worse as you suddenly find some teams have more data than others, and they use features differently, so you either need a ton of alarms or fewer messier ones, but you're getting further and further away from whether people are actually satisfied.
This loss of context is a critical part of dealing with systems that are too complex to adequately be represented by a single aggregate. Whenever a signal is useful, an in-depth dive is usually worth it if you are looking to embrace complexity.
The metric is better used to attract your attention than as a target or as something that tells you what to know. Seek to explain and understand the metric first, not to change it.
![A OSEFW INDICAToR ITSELF USELESS (i ROhJSTCY RAVE GØD Foe A t.JlCELY ](Exported%20image%2020240808113916-17.png)
As a related concept, if you act on a leading indicator, it stops leading, particularly when its influenced by trade-offs.
Metrics that become their own targets and are gamed of course lose meaningfulness; this is one of the most common issues with counting incidents and then debating whether an outage should or shouldnt be declared in a way that might affect the tally rather than addressing it directly.
But other metrics are of interest as well. If you evaluate your total capacity by some bottlenecks value, and that this bottleneck is a target of optimization work, you will lose the ability to easily know when or how to scale up because that bottleneck possibly hid something else. This is contributing to a non-negligible portion of our incidents at work I believe. We fix a thing that acted as an implicit blocker and off we go into the great unknown.
Our storage engine's disk storage used to be our main bottleneck. We drove scaling out and rebalancing traffic based on how close we were to heavy usage across multiple partitions. This was a useful signal, but it also drove costs up, and eventually became the target of optimization.
An engineer successfully made our data offloading almost an order of magnitude faster, and eliminated our most glaring scaling issues at the time. Removing this limit however messed with our ability to know when to scale, which then revealed issues with file descriptors, memory, and snapshotting times.
The only good advice I have here is to re-evaluate your metrics often, and change them. I guess theres also a lesson to be learned that improvements can also cause their own uncertainty and that these successes can themselves lead to destabilizations.
Because we no longer needed to scale out as aggressively and were free to discover new issues, and one of our best improvements to the system in recent memory is therefore also a contributor to a lot of operational challenges.
![PEOPLE WILL Do WRAT THEY BELIEVE IS USEFUL ](Exported%20image%2020240808113916-18.png)
Things that people think are useful are possibly going to happen even if you forbid them. If you forbid people from logging onto production hosts, and they truly think they'll need it for emergency situations, they'll make sure there's still a way for it to happen, albeit under a different name.
On the other hand, things that people think are useless are likely to be done in a minimal way with no enthusiasm, such as lying in your timesheets.
This means that writing a procedure means little unless people actually see its value and believe its worth following. Conversely, it means that if you can demonstrate the usefulness and make some approaches more usable, theyre likely to get adopted regardless of what is written down as a list of steps or procedures.
A related concept here is one here is that if you are tracking things like action items after an incident reviews and they go in the backlog to die, it may not be that your people are failing to follow through; it might also be that its impractical to do so, or its could also be that these action items were never feeling useful, and the process itself needs to be revisited rather than reinforced.
Seeing non-compliance is not necessarily a sign of bad workers. It may rather be a sign of a bad understanding of the workers' challenges, and point to a need to adjust how work is prescribed.
Getting a small real buy-in into something voluntary may be better than getting fake buy-in into something youre forcing people to do. Of course if you manage to write a good procedure that people believe are worth following, more power to you, this is going great.
![A TEAM Yao.) RANDN PEOPLE ](Exported%20image%2020240808113916-19.png)
The shortest feedback loop may be attained by giving people the tools to make the right decisions right there and then, and let them do it. Cut the middlemen, including yourself.
How do you make that work? We come back to goal alignments and top priorities being harmonized and well understood. If the pressures and goals are understood better, the decisions made also work better.
That does mean that you have to listen back about how these things have been going, and that not only do you need to trust your people, but they need to trust you back with critical and unpleasant information as well. The feedback flows both ways, and this hinges on psychological safety.
If you've ever talked to a contractor asked to help a big organization, the first thing they'll tell you they do is go talk to the workers with boots on the ground, and ask them what they think needs changing. They'll often have years of potential improvements backlogged, and that they're ready to tell anyone about. Either because management wouldn't listen to it, or because the workers lost trust that voicing that feedback would yield any result.
Then the contractor brings it up to management as a neutral party, and suddenly it gets listened to and acted upon.
If you've lost that trust, then contractors can play that specific role of workers at the periphery of the organization helping drive change, and they can play a very useful function.
But if you have that trust already, maintaining it is crucial because thats how you get all the good information to help orient and influence things.
Trust also means that if you want people to be innovative, you have to allow them to make mistakes. You cant get it right the first time all the time; if people cant be allowed to get it wrong here and there, they wont be allowed to improve and try new things either.
![Mow To EMBRACE COMPLEXITY? QADE-OFFs FVk1kJG KEEP EEDBACk LOPS TIGHT HAC— — H OVER ](Exported%20image%2020240808113916-20.png)
Finally, let's look at shifting perspective away from a bare analysis and onto a more systemic point of view. People in specific teams often have a more detailed expert view than you could either have, but if you're standing outside of it, your strength might be to understand how the parts interact in a way that isn't visible to the inside.
![ITS HARD WITHOUT CHANGING TE PRESSES MT @STEe ](Exported%20image%2020240808113916-21.png)
The most basic point here is that you cant expect to change the outcome of these small little decisions that accumulate all the time if you never address the pressures within the system that foster them.
I used to try and weed my lawn a whole hell of a lot and pull the weeds hours a week until someone explained to me that weeds grew easier in the type of soil I had (poor, dry, unmaintained soil) than grass, and pulling the weeds wasnt the way to go, I needed to actually make the soil good for the grass to crowd out the weeds.
It's similar when considering this whole idea of root cause analysis—of trying to find the one source of the problem and removing it. If your root cause is at the weeds level, youll keep pulling on them forever and will rarely make decent progress. The weeds will keep growing no matter how many roots you remove.
If you foster good soil, if you create the right environment that encourages the type of behavior you want instead of the type of behaviour you dislike, you have hopes that the good stuff will crowd out the bad stuff. Thats a roundabout way of talking about culture change. And for these, deep dives based on [richer narratives](https://ferd.ca/notes/paper-accident-report-interpretation.html) and [thematic analysis](https://www.jeli.io/howie/welcome) prove more useful.
Also there's a warning here about trying to change the decisions your people make with carrots and sticks—with incentives. They are not going to fundamentally change what pressures the employees negotiate. The pressures stay the same, all you're doing is adding more of them, either in the form of rewards or punishments, which makes decision-making more complex and trickier.
Chances are people will keep making the same decisions as they were already, but then they'll report it differently to either get their bonus or to avoid getting penalized for it. Surfacing, understanding, and clarifying goal conflicts can make things easier or shape work to give them more room. Adding carrots and sticks can make things harder.
![ITS HARD TO WITHOUT CHANGING TE PRESS)ES MT @STEe ](Exported%20image%2020240808113916-22.png)
But the tip here is probably: look into what are the behaviors you want to see happen, and give them room to grow.
My most successful initiative at Honeycomb is probably creating [weekly discussion sessions about operational stuff and on-call](https://www.honeycomb.io/blog/oncallogy-sessions-best-practices). They range from “how do we operate new service X” into trickier discussions like “is it okay to be visibly angry in an incident”, “how do you deal with shit you dont know or avoid burnout” or “are there times where code freezes are actually a useful thing?”.
Over time we looked into all sorts of weird interactions and the meeting became its own tool.
When we noticed incident reviews were difficult to schedule across departments and timezones, we decided that a good wide incident review is good operational talk and started making the optional time slot, which was already on every engineer's calendar (and some other departments too), available for them. It became easier for people to run incident reviews, and over time their size grew from 7-8 people, scoped to 1 or 2 teams, to bigger events with 20 to 40 people in them.
We removed a huge but subtle blocker to good feedback loops existing within the organization.
These sorts of small changes are those you can drive locally with almost no risk of having them run afoul of organizational priorities, and when you see them work, use the org structure to expand them everywhere.
![IODICATOR IS VSEFUL WH24 IS AcTED ON -12. 70.710/ o 94.13% ](Exported%20image%2020240808113916-23.png)
I find it useful to keep focusing on what an indicator triggers as a behavior (the interaction) rather than _only_ what it reports directly. This slide here is 4 error budgets from our SLOs, which combine how successful requests are both in terms of speed and errors, compared to an objective we express in terms of the desired fault rate.
When we have to pick targets for our platform, people often ask whether we could pick some key SLOs and turn them as the objective. My answer is almost always "I don't care if we meet the SLOs or not". I mean I care, but not like that.
SLOs arent hard and fast rules. When the error budget is empty, the main thing that matters to me is that we have a conversation about it, and decide what it is we want to happen from there on. Are we going to hold off on deploys and experiments? Are we able to meet the objectives while on-call, with some schedule corrective work, some major re-architecting? Can we just talk to the customers? Were our targets too ambitious or are we going to eat dirt for a while?
Kneejerk automated reactions arent nearly as useful as sitting down and having a cross-departmental discussion about what it is we want to do, as an organization, about these signals of unmet expectations. If it fits within on-call duty, like what is probably the case with the error budget on the top left, then fine.
But in other cases, such as the top right budget here, which seems to show a gradual decline, owe have to choose whether to do corrective work (and how/when) to meet the SLO—because that wasn't expected and is undesirable—or maybe to relax it—because that's actually a natural consequence of new more expensive features and we need to tweak definitions. Or we could temporarily ignore it because corrective work is already on the way, but not a top priority right now.
The two budgets at the bottom come from SLOs that may never page anyone. But from time to time, we re-calibrate them by asking support whether there are any issues users complain about that we aren't already aware of. So long as we're ahead of the complaints, we figure the SLOs are properly defined. But from time to time, we find out that we slipped by getting comments on things our alerting never properly captured. Or maybe we needed to better manage the user's expectations—that's also an option.
For any of these choices, we also have to know how this is going to be communicated to users and customers, and having these discussions is the true value of SLOs to me. SLOs that flow outside of engineering teams provide a greater feedback loop about our practices, further upstream, than those that are used exclusively by the teams defining them, regardless of their use for alerting.
![BE THE FEEDBACK Loop ](Exported%20image%2020240808113916-24.png)
Finally, this is where SREs can be placed in a great way to shine. You can be away from the central roles, away from the decision-making, on the periphery. By being outside of silos and floating around the organizations structure, you are allowed to take information from many levels, carry it around, and really tie the loop at the end of so many decisions made in the organization by noting and carrying their impact back once theyve hit a production system.
It is an iterative exercise, our sociotechnical systems are alive, and carrying pertinent signals and amplifying them, you can influence how long its gonna take before it all goes to hell anyway.

View File

@@ -0,0 +1,43 @@
Clipped from: [https://www.primevideotech.com/video-streaming/scaling-up-the-prime-video-audio-video-monitoring-service-and-reducing-costs-by-90](https://www.primevideotech.com/video-streaming/scaling-up-the-prime-video-audio-video-monitoring-service-and-reducing-costs-by-90)
## The move from a distributed microservices architecture to a monolith application helped achieve higher scale, resilience, and reduce costs.
At Prime Video, we offer thousands of live streams to our customers. To ensure that customers seamlessly receive content, Prime Video set up a tool to monitor every stream viewed by customers. This tool allows us to automatically identify perceptual quality issues (for example, block corruption or audio/video sync problems) and trigger a process to fix them.
Our Video Quality Analysis (VQA) team at Prime Video already owned a tool for audio/video quality inspection, but we never intended nor designed it to run at high scale (our target was to monitor thousands of concurrent streams and grow that number over time). While onboarding more streams to the service, we noticed that running the infrastructure at a high scale was very expensive. We also noticed scaling bottlenecks that prevented us from monitoring thousands of streams. So, we took a step back and revisited the architecture of the existing service, focusing on the cost and scaling bottlenecks.
The initial version of our service consisted of distributed components that were orchestrated by [AWS Step Functions](https://docs.aws.amazon.com/step-functions/latest/dg/welcome.html). The two most expensive operations in terms of cost were the orchestration workflow and when data passed between distributed components. To address this, we moved all components into a single process to keep the data transfer within the process memory, which also simplified the orchestration logic. Because we compiled all the operations into a single process, we could rely on scalable [Amazon Elastic Compute Cloud (Amazon EC2)](https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html) and [Amazon Elastic Container Service (Amazon ECS)](https://docs.aws.amazon.com/AmazonECS/latest/developerguide/Welcome.html) instances for the deployment.
### **Distributed systems overhead**
Our service consists of three major components. The media converter converts input audio/video streams to frames or decrypted audio buffers that are sent to detectors. Defect detectors execute algorithms that analyze frames and audio buffers in real-time looking for defects (such as video freeze, block corruption, or audio/video synchronization problems) and send real-time notifications whenever a defect is found. For more information about this topic, see our [How Prime Video uses machine learning to ensure video quality](https://www.primevideotech.com/computer-vision/how-prime-video-uses-machine-learning-to-ensure-video-quality) article. The third component provides orchestration that controls the flow in the service.
We designed our initial solution as a distributed system using serverless components (for example, AWS Step Functions or [AWS Lambda](https://docs.aws.amazon.com/lambda/latest/dg/welcome.html)), which was a good choice for building the service quickly. In theory, this would allow us to scale each service component independently. However, the way we used some components caused us to hit a hard scaling limit at around 5% of the expected load. Also, the overall cost of all the building blocks was too high to accept the solution at a large scale.
The following diagram shows the serverless architecture of our service.
![The diagram shows a control plane and data plan in the initial architecture. The customer's request is handled by a lambda function that is then forwarded to relevant step functions that execute detectors. At the same time, Media Conversion service starts processing the input stream, providing artifacts to detectors through an S3 bucket. Once the analysis is completed, the aggregated result is being stored in an S3 bucket.](Exported%20image%2020240808113914-0.png)
**The initial architecture of our defect detection system.**
The main scaling bottleneck in the architecture was the orchestration management that was implemented using AWS Step Functions. Our service performed multiple state transitions for every second of the stream, so we quickly reached account limits. Besides that, AWS Step Functions charges users per state transition.
The second cost problem we discovered was about the way we were passing video frames (images) around different components. To reduce computationally expensive video conversion jobs, we built a microservice that splits videos into frames and temporarily uploads images to an [Amazon Simple Storage Service (Amazon S3)](https://docs.aws.amazon.com/AmazonS3/latest/userguide/Welcome.html) bucket. Defect detectors (where each of them also runs as a separate microservice) then download images and processed it concurrently using AWS Lambda. However, the high number of Tier-1 calls to the S3 bucket was expensive.
### **From distributed microservices to a monolith application**
To address the bottlenecks, we initially considered fixing problems separately to reduce cost and increase scaling capabilities. We experimented and took a bold decision: we decided to rearchitect our infrastructure.
We realized that distributed approach wasnt bringing a lot of benefits in our specific use case, so we packed all of the components into a single process. This eliminated the need for the S3 bucket as the intermediate storage for video frames because our data transfer now happened in the memory. We also implemented orchestration that controls components within a single instance.
The following diagram shows the architecture of the system after migrating to the monolith.
![The diagram represents a control and data plan for the updated architecture. All the components run within a single ECS task, therefore the control doesn't go through the network. Data sharing is done through instance memory and only the final results are uploaded to an S3 bucket.](Exported%20image%2020240808113914-1.png)
**The updated architecture for monitoring a system with all components running inside a single Amazon ECS task.**
Conceptually, the high-level architecture remained the same. We still have exactly the same components as we had in the initial design (media conversion, detectors, or orchestration). This allowed us to reuse a lot of code and quickly migrate to a new architecture.
In the initial design, we could scale several detectors horizontally, as each of them ran as a separate microservice (so adding a new detector required creating a new microservice and plug it in to the orchestration). However, in our new approach the number of detectors only scale vertically because they all run within the same instance. Our team regularly adds more detectors to the service and we already exceeded the capacity of a single instance. To overcome this problem, we cloned the service multiple times, parametrizing each copy with a different subset of detectors. We also implemented a lightweight orchestration layer to distribute customer requests.
The following diagram shows our solution for deploying detectors when the capacity of a single instance is exceeded.
![Customer's request is being forwarded by a lambda function to relevant ECS tasks. The result for each detector is stored in S3 bucket separately.](Exported%20image%2020240808113914-2.png)
**Our approach for deploying more detectors to the service.**
### **Results and takeaways**
Microservices and serverless components are tools that do work at high scale, but whether to use them over monolith has to be made on a case-by-case basis.
Moving our service to a monolith reduced our infrastructure cost by over 90%. It also increased our scaling capabilities. Today, were able to handle thousands of streams and we still have capacity to scale the service even further. Moving the solution to Amazon EC2 and Amazon ECS also allowed us to use the [Amazon EC2 compute saving plans](https://aws.amazon.com/savingsplans/compute-pricing/) that will help drive costs down even further.
Some decisions weve taken are not obvious but they resulted in significant improvements. For example, we replicated a computationally expensive media conversion process and placed it closer to the detectors. Whereas running media conversion once and caching its outcome might be considered to be a cheaper option, we found this not be a cost-effective approach.
The changes weve made allow Prime Video to monitor all streams viewed by our customers and not just the ones with the highest number of viewers. This approach results in even higher quality and an even better customer experience.

View File

@@ -0,0 +1,21 @@
Clipped from: [https://blog.traillifeusa.com/boy-finds-hope?utm_campaign=Raising%20Godly%20Boys&utm_medium=email&_hsmi=220415115&_hsenc=p2ANqtz-9wOKKcP8ywC8lxPGr9wAs3lzk6NsuStGaUd8Jw9bL9FQtvKma9aP_6LMzlBiba-jFHNF60TmvQ-95DT6tkKo_m8gvHig&utm_content=220415115&utm_source=hs_automation](https://blog.traillifeusa.com/boy-finds-hope?utm_campaign=Raising%20Godly%20Boys&utm_medium=email&_hsmi=220415115&_hsenc=p2ANqtz-9wOKKcP8ywC8lxPGr9wAs3lzk6NsuStGaUd8Jw9bL9FQtvKma9aP_6LMzlBiba-jFHNF60TmvQ-95DT6tkKo_m8gvHig&utm_content=220415115&utm_source=hs_automation)
![Struggling Boy Finds Hope, Purpose, and Self-Worth through Trail Life Mentors](Exported%20image%2020240808113917-0%201.jpeg)
**Trail Life mentors inspire young boy not only to find the will to live, but to help others around him**
 
In a society where the lines between masculine and feminine are constantly blurred, brash boyish bluster and boisterousness is losing its place. Feeling unappreciated, too boys are losing their identity. Traditional needs like physical challenge, competition, risk-taking, action, and adventure are being discounted and boys seeking to fill these needs are punished, diagnosed, or written off as unruly, difficult, or perhaps toxic.
Trail Life USA, the largest Christ-centered, boy-focused scout-type organization in the country, is familiar with this struggle for boys and can provide a solution. In the fight to provide boys with unique programming that celebrates boyhood, Trail Life utilizes outdoor adventure and personal relationships to speak to the heart of a boy and and to guide him in his walk with Christ.
Mark Hancock, Trail Life CEO, commented, _“When boys feel like they are relegated as less than due to unappreciated gender differences, they begin to wonder where they belong in a society that seems to discount their abilities. Trail Life USA provides a boy-focused program and activities designed to let boys be boys, accentuating their strengths and allowing them to feel understood and appreciated.”_
_Hancock continued, “It seems everywhere a boy goes, hes expected to comply with unrealistic social norms. The consistent message he gets is that he needs to sit still, be quiet, do what he is told, and behave like the girls. But boys are not defective girls. They are created differently on purpose for a purpose. Properly channeled and intentionally challenged, the exuberance, drive, and daring of healthy boys is exactly what our society needs.”_
[![New call-to-action](Exported%20image%2020240808113917-1.jpeg)](https://blog.traillifeusa.com/cs/c/?cta_guid=ae0cca4b-12ee-44b5-9522-ea34bcd24e9e&signature=AAH58kHdz_P7HD13orXjY97P9tTqTGFjMw&pageId=77377023317&placement_guid=fe90369d-b6b8-4f68-9fc0-b6d63f14044c&click=19d2745d-b484-4784-8cde-ecde604c2e2a&hsutk=&canon=https%3A%2F%2Fblog.traillifeusa.com%2Fboy-finds-hope&portal_id=6459804&redirect_url=APefjpGnipquSbf2uJG8qs097FtoxGreH-0PhQnuOewllWKAO8ER1syjqeWtpneVenoGaHglMjK-_DfDkJUb2mKP7Sb3d5-H0WCHgYDmgDyzI-ZHx59iF80L-97IR8vDn99tCMt7ezx2Rj6p4CLK-nYdn7aZHiWVcw)
In a system that fails to acknowledge that boys are not just like girls, [boys are increasingly diagnosed with disorders](https://www.understood.org/articles/en/do-boys-have-learning-and-thinking-differences-more-often-than-girls) and are [falling behind their female counterparts](https://www.brookings.edu/blog/up-front/2021/01/12/the-unreported-gender-gap-in-high-school-graduation-rates/) in nearly every academic category. This trend continues into [college where 60% of students are female](https://www.usatoday.com/story/opinion/2021/10/09/boys-falling-behind-how-schools-must-change-help-young-males/5913463001/). Even more tragic than the academic difficulty is the impact on the mental health of boys and men. Today, men account for [four out of five suicides in America](https://www.cdc.gov/nchs/products/databriefs/db373.htm#:~:text=In%2520both%2520urban%2520and%2520rural%2520areas%252C%2520suicide%2520rates%2520for%2520males,(30.7%2520compared%2520with%25208.0).) and twice the [drug-related deaths](https://news.wttw.com/2021/10/25/us-overdose-deaths-surge-all-time-high) as compared to women. The most rapidly growing suicide rate demographic is boys from the [ages of ten to 14](https://www.bloomberg.com/news/articles/2021-11-03/u-s-suicides-fall-for-second-year-in-a-row-during-pandemic). One mother recently wrote to share how her son was almost one of these statistics:
_“Dealing with ADHD, depression, and anxiety, my son was struggling at school and at home. Daily outbursts, disciplinary problems, and panic attacks forced us as parents to make the hard decision to pull him out of traditional school. A constant cycle of being in trouble with teachers, church leaders, and parents left him questioning his own value. At ten years of age, he decided that he was so broken and hopeless that the world would be better off without him. Last October, he attempted to take his own life and was admitted to the emergency room at the childrens hospital because of injuries sustained in that attempt_.
_“Then my son began attending_ _Trail Life__. It was the first time that he felt understood and accepted by authority figures (the_ **Trail Life** _leaders). In a world that had always attempted to squash his character, he has been encouraged to see a purpose in his boy-ness. He has heard consistently that God created him the way he is and that he is loved. He has been encouraged that the future holds a purpose for him — one that will utilize his courage, his sensitivity, and his passion. He has been inspired intellectually and spiritually. Most importantly, he feels part of a community where he respects and admires the leaders, and has friends.”_ 
In the active learning environment at **Trail Life**, her son has been able to shine. In the past six months, he has been hiking and camping, learned first aid, become proficient in lighting a fire without a match, and built rockets with his Troop. 
The mother commented on these skills her son learned at Trail Life, stating, _“Because my son learned first aid with his Troop, I still have both of my children. Last Sunday, I was driving on the freeway when my daughter, who is two years old, choked on a snack. She couldnt breathe. I was stuck driving in traffic in the HOV lane, where I couldnt pull over. My son was able to use the skills he learned at Trail Life_ _to give his sister the Heimlich maneuver, and she coughed up food and was able to breathe. He saved her life and displayed the level-headed, calm, and confident skills he needed to save his sister. I am so proud of him, and so thankful for Trail Life."_
_“Even his father is seeing a tremendous difference. Before my sons attempted suicide, they had become so estranged they were barely able to talk. When he was presented with the Life Saving Award from Trail Life,_ _his dad was able to attend and publicly commend his son. The young man who had been labeled as troublesome, difficult, and delinquent by so many is now being commended, affirmed, accepted, and encouraged by his Troop, his community, and his father.”_
_“Many sons never hear words of affirmation, acceptance, and encouragement from their father. I want to thank the men of the Troop who have poured themselves into my son and our family. Because of your intervention and effort, a relationship that was broken has been dramatically healed and turned around.”_
The mother concluded, _“Trail Life is m__aking a real difference in the lives of boys and their fathers.”_

View File

@@ -0,0 +1,13 @@
Clipped from: [https://blog.traillifeusa.com/what-is-a-boy?utm_campaign=Raising%20Godly%20Boys&utm_medium=email&_hsmi=219419808&_hsenc=p2ANqtz-9tkiaPng4HKJvYRy6Ee8wBSX_GRpZ-nKWnj4tZJxJ13W0R00aCDx3IevUfMoI2iUWmB4GMzcGgFwHTSSQIWosWfF_FPw&utm_content=219419808&utm_source=hs_automation](https://blog.traillifeusa.com/what-is-a-boy?utm_campaign=Raising%20Godly%20Boys&utm_medium=email&_hsmi=219419808&_hsenc=p2ANqtz-9tkiaPng4HKJvYRy6Ee8wBSX_GRpZ-nKWnj4tZJxJ13W0R00aCDx3IevUfMoI2iUWmB4GMzcGgFwHTSSQIWosWfF_FPw&utm_content=219419808&utm_source=hs_automation)
![What is a Boy?](Exported%20image%2020240808113917-0%201.jpeg)
Between the innocence of babyhood and the dignity of manhood we find a delightful creature called a boy. Boys come in assorted sizes, weights, and colors, but all boys have the same creed: to enjoy every second of every minute of every hour of every day and to protest with noise (their only weapon) when their last minute is finished and the adult males pack them off to bed at night.
Boys are found everywhere—on top of, underneath, inside of, climbing on, swinging from, running around, or jumping to.
Mothers love them, little girls hate them, older sisters and brothers tolerate them, adults ignore them, and Heaven protects them.
A boy is truth with dirt on its face, beauty with a cut on its finger, wisdom with bubble gum in its hair, and the hope of the future with a frog in its pocket. When you are busy, a boy is an inconsiderate, bothersome, intruding jangle of noise. When you want him to make a good impression, his brain turns to jelly or else he becomes a savage, sadistic, jungle creature bent on destroying the world and himself with it.
A boy is a composite—he has the appetite of a horse, the digestion of a sword-swallower, the energy of a pocket-sized atomic bomb, the curiosity of a cat, the lungs of a dictator, the imagination of a Paul Bunyan, the shyness of a violet, the audacity of a steel trap, the enthusiasm of a firecracker, and when he makes something, he has five thumbs on each hand. He likes ice cream, knives, saws, Christmas, comic books, the boy across the street, woods, water (in its natural habitat), large animals, Dad, trains, Saturday mornings, and fire engines.
He is not much for Sunday School, company, schools, books without pictures, music lessons, neckties, barbers, girls, overcoats, adults, or bedtime. Nobody else is so early to rise, or so late to supper. Nobody else gets so much fun out of trees, dogs, and breezes. Nobody else can cram into one pocket a rusty knife, a half-eaten apple, three feet of string, an empty Bull Durham sack, two gum drops, six cents, a slingshot, a chunk of unknown substance, and a genuine supersonic code ring with a secret compartment.
A boy is a magical creature—you can lock him out of your workshop, but you cant lock him out of your heart. You can get him out of your study, but you cant get him out of your mind. Might as well give up—he is your captor, your jailer, your boss, and your master—a freckled-faced, pint-sized, cat-chasing, bundle of noise. But when you come home at night with only shattered pieces of your hopes and dreams, he can mend them like new with two magic words, "Hi Dad!"
---
[Trail Life USA](http://www.traillifeusa.com/) is designed uniquely for boys. Established on timeless values and set in the context of outdoor adventure, boys from Kindergarten through 12th grade are engaged in a Troop setting by male mentors where they are challenged to grow in character, understand their purpose, serve their community, and develop practical leadership skills to carry out the mission for which they were created.

View File

@@ -0,0 +1,3 @@
| | |
|---|---|
|![thumbnail](Exported%20image%2020240808113925-0.png)|\| \|<br>\|---\|<br>\|## 2020-01-31 Tech Talk - Michael Sterling - Royalties - Faithlife Coders - Amber\|<br>\|[https://amber.faithlife.com/shares/921WWa1mhUUOlewF](https://amber.faithlife.com/shares/921WWa1mhUUOlewF)\|<br>\|Royalties Tech Talk Slides: [https://docs.google.com/presentation/d/1UERbx5Op1upek32_TCBuXRPgIKtZkXjmWQM6IyEzCSA/edit#slide=id.g7d0408af6d_2_93](https://docs.google.com/presentation/d/1UERbx5Op1upek32_TCBuXRPgIKtZkXjmWQM6IyEzCSA/edit#slide=id.g7d0408af6d_2_93)...\||

View File

@@ -0,0 +1,415 @@
Clipped from: [https://applied-llms.org/](https://applied-llms.org/)
A practical guide to building successful LLM products.
Authors
[Eugene Yan](https://eugeneyan.com/)
[Bryan Bischof](https://www.linkedin.com/in/bryan-bischof/)
[Charles Frye](https://www.linkedin.com/in/charles-frye-38654abb/)
[Hamel Husain](https://hamel.dev/)
[Jason Liu](https://jxnl.co/)
[Shreya Shankar](https://www.sh-reya.com/)
Published
June 8, 2024
Also published on OReilly Media in three parts: [Tactical](https://www.oreilly.com/radar/what-we-learned-from-a-year-of-building-with-llms-part-i/), [Operational](https://www.oreilly.com/radar/what-we-learned-from-a-year-of-building-with-llms-part-ii/), Strategic (pending).
Its an exciting time to build with large language models (LLMs). Over the past year, LLMs have become “good enough” for real-world applications. And theyre getting better and cheaper every year. Coupled with a parade of demos on social media, there will be an [estimated $200B investment in AI by 2025](https://www.goldmansachs.com/intelligence/pages/ai-investment-forecast-to-approach-200-billion-globally-by-2025.html). Furthermore, provider APIs have made LLMs more accessible, allowing everyone, not just ML engineers and scientists, to build intelligence into their products. Nonetheless, while the barrier to entry for building with AI has been lowered, creating products and systems that are effective—beyond a demo—remains deceptively difficult.
Weve spent the past year building, and have discovered many sharp edges along the way. While we dont claim to speak for the entire industry, wed like to share what weve learned to help you avoid our mistakes and iterate faster. These are organized into three sections:
- [Tactical](https://applied-llms.org/#tactical-nuts--bolts-of-working-with-llms): Some practices for prompting, RAG, flow engineering, evals, and monitoring. Whether youre a practitioner building with LLMs, or hacking on weekend projects, this section was written for you.
- [Operational](https://applied-llms.org/#operation-day-to-day-and-org-concerns): The organizational, day-to-day concerns of shipping products, and how to build an effective team. For product/technical leaders looking to deploy sustainably and reliably.
- Strategic: The long-term, big-picture view, with opinionated takes such as “no GPU before PMF” and “focus on the system not the model”, and how to iterate. Written with founders and executives in mind.
We intend to make this a practical guide to building successful products with LLMs, drawing from our own experiences and pointing to examples from around the industry.
Ready to ~~delve~~ dive in? Lets go.
# 1 Tactical: Nuts & bolts of working with LLMs
Here, we share best practices for core components of the emerging LLM stack: prompting tips to improve quality and reliability, evaluation strategies to assess output, retrieval-augmented generation ideas to improve grounding, how to design human-in-the-loop workflows, and more. While the technology is still nascent, we trust these lessons are broadly applicable and can help you ship robust LLM applications.
## 1.1 Prompting
We recommend starting with prompting when prototyping new applications. Its easy to both underestimate and overestimate its importance. Its underestimated because the right prompting techniques, when used correctly, can get us very far. Its overestimated because even prompt-based applications require significant engineering around the prompt to work well.
### 1.1.1 Focus on getting the most out of fundamental prompting techniques
A few prompting techniques have consistently helped with improving performance across a variety of models and tasks: n-shot prompts + in-context learning, chain-of-thought, and providing relevant resources.
The idea of in-context learning via n-shot prompts is to provide the LLM with examples that demonstrate the task and align outputs to our expectations. A few tips:
- If n is too low, the model may over-anchor on those specific examples, hurting its ability to generalize. As a rule of thumb, aim for n ≥ 5. Dont be afraid to go as high as a few dozen.
- Examples should be representative of the prod distribution. If youre building a movie summarizer, include samples from different genres in roughly the same proportion youd expect to see in practice.
- You dont always need to provide the input-output pairs; examples of desired outputs may be sufficient.
- If you plan for the LLM to use tools, include examples of using those tools.
In Chain-of-Thought (CoT) prompting, we encourage the LLM to explain its thought process before returning the final answer. Think of it as providing the LLM with a sketchpad so it doesnt have to do it all in memory. The original approach was to simply add the phrase “Lets think step by step” as part of the instructions, but, weve found it helpful to make the CoT more specific, where adding specificity via an extra sentence or two often reduces hallucination rates significantly.
For example, when asking an LLM to summarize a meeting transcript, we can be explicit about the steps:
- First, list out the key decisions, follow-up items, and associated owners in a sketchpad.
- Then, check that the details in the sketchpad are factually consistent with the transcript.
- Finally, synthesize the key points into a concise summary.
Note that in recent times, [some doubt](https://arxiv.org/abs/2405.04776) has been cast on if this technique is as powerful as believed. Additionally, theres significant debate as to exactly what is going on during inference when Chain-of-Thought is being used. Regardless, this technique is one to experiment with when possible.
Providing relevant resources is a powerful mechanism to expand the models knowledge base, reduce hallucinations, and increase the users trust. Often accomplished via Retrieval Augmented Generation (RAG), providing the model with snippets of text that it can directly utilize in its response is an essential technique. When providing the relevant resources, its not enough to merely include them; dont forget to tell the model to prioritize their use, refer to them directly, and to mention when none of the resources are sufficient. These help “ground” agent responses to a corpus of resources.
### 1.1.2 Structure your inputs and outputs
Structured input and output help models better understand the input as well as return output that can reliably integrate with downstream systems. Adding serialization formatting to your inputs can help provide more clues to the model as to the relationships between tokens in the context, additional metadata to specific tokens (like types), or relate the request to similar examples in the models training data.
As an example, many questions on the internet about writing SQL begin by specifying the SQL schema. Thus, you can expect that effective prompting for Text-to-SQL should include [structured schema definitions](https://www.researchgate.net/publication/371223615_SQL-PaLM_Improved_Large_Language_ModelAdaptation_for_Text-to-SQL).
Structured input expresses tasks clearly and resembles how the training data is formatted, increasing the probability of better output. Structured output simplifies integration into downstream components of your system. [Instructor](https://github.com/jxnl/instructor) and [Outlines](https://github.com/outlines-dev/outlines) work well for structured output. (If youre importing an LLM API SDK, use Instructor; if youre importing Huggingface for a self-hosted model, use Outlines.)
When using structured input, be aware that each LLM family has their own preferences. Claude prefers <xml> while GPT favors Markdown and JSON. With XML, you can even pre-fill Claudes responses by providing a <response> tag like so.
messages=[ { "role": "user", "content": """Extract the <name>, <size>, <price>, and <color> from this product description into your <response>. <description>The SmartHome Mini is a compact smart home assistant available in black or white for only $49.99. At just 5 inches wide, it lets you control lights, thermostats, and other connected devices via voice or app—no matter where you place it in your home. This affordable little hub brings convenient hands-free control to your smart devices. </description>""" }, { "role": "assistant", "content": "<response><name>" }]
### 1.1.3 Have small prompts that do one thing, and only one thing, well
A common anti-pattern / code smell in software is the “[God Object](https://en.wikipedia.org/wiki/God_object)”, where we have a single class or function that does everything. The same applies to prompts too.
A prompt typically starts simple: A few sentences of instruction, a couple of examples, and were good to go. But as we try to improve performance and handle more edge cases, complexity creeps in. More instructions. Multi-step reasoning. Dozens of examples. Before we know it, our initially simple prompt is now a 2,000 token Frankenstein. And to add injury to insult, it has worse performance on the more common and straightforward inputs! GoDaddy shared this challenge as their [No. 1 lesson from building with LLMs](https://www.godaddy.com/resources/news/llm-from-the-trenches-10-lessons-learned-operationalizing-models-at-godaddy#h-1-sometimes-one-prompt-isn-t-enough).
Just like how we strive (read: struggle) to keep our systems and code simple, so should we for our prompts. Instead of having a single, catch-all prompt for the meeting transcript summarizer, we can break it into steps:
- Extract key decisions, action items, and owners into structured format
- Check extracted details against the original transcription for consistency
- Generate a concise summary from the structured details
As a result, weve split our single prompt into multiple prompts that are each simple, focused, and easy to understand. And by breaking them up, we can now iterate and eval each prompt individually.
### 1.1.4 Craft your context tokens
Rethink, and challenge your assumptions about how much context you actually need to send to the agent. Be like Michaelangelo, do not build up your context sculpture—chisel away the superfluous material until the sculpture is revealed. RAG is a popular way to collate all of the potentially relevant blocks of marble, but what are you doing to extract whats necessary?
Weve found that taking the final prompt sent to the model—with all of the context construction, and meta-prompting, and RAG results—putting it on a blank page and just reading it, really helps you rethink your context. We have found redundancy, self-contradictory language, and poor formatting using this method.
The other key optimization is the structure of your context. If your bag-of-docs representation isnt helpful for humans, dont assume its any good for agents. Think carefully about how you structure your context to underscore the relationships between parts of it and make extraction as simple as possible.
More [prompting fundamentals](https://eugeneyan.com/writing/prompting/) such as prompting mental model, prefilling, context placement, etc.
## 1.2 Information Retrieval / RAG
Beyond prompting, another effective way to steer an LLM is by providing knowledge as part of the prompt. This grounds the LLM on the provided context which is then used for in-context learning. This is known as retrieval-augmented generation (RAG). Practitioners have found RAG effective at providing knowledge and improving output, while requiring far less effort and cost compared to finetuning.
### 1.2.1 RAG is only as good as the retrieved documents relevance, density, and detail
The quality of your RAGs output is dependent on the quality of retrieved documents, which in turn can be considered along a few factors
The first and most obvious metric is relevance. This is typically quantified via ranking metrics such as [Mean Reciprocal Rank (MRR)](https://en.wikipedia.org/wiki/Mean_reciprocal_rank) or [Normalized Discounted Cumulative Gain (NDCG)](https://en.wikipedia.org/wiki/Discounted_cumulative_gain). MRR evaluates how well a system places the first relevant result in a ranked list while NDCG considers the relevance of all the results and their positions. They measure how good the system is at ranking relevant documents higher and irrelevant documents lower. For example, if were retrieving user summaries to generate movie review summaries, well want to rank reviews for the specific movie higher while excluding reviews for other movies.
Like traditional recommendation systems, the rank of retrieved items will have a significant impact on how the LLM performs on downstream tasks. To measure the impact, run a RAG-based task but with the retrieved items shuffled—how does the RAG output perform?
Second, we also want to consider information density. If two documents are equally relevant, we should prefer one thats more concise and has fewer extraneous details. Returning to our movie example, we might consider the movie transcript and all user reviews to be relevant in a broad sense. Nonetheless, the top-rated reviews and editorial reviews will likely be more dense in information.
Finally, consider the level of detail provided in the document. Imagine were building a RAG system to generate SQL queries from natural language. We could simply provide table schemas with column names as context. But, what if we include column descriptions and some representative values? The additional detail could help the LLM better understand the semantics of the table and thus generate more correct SQL.
### 1.2.2 Dont forget keyword search; use it as a baseline and in hybrid search
Given how prevalent the embedding-based RAG demo is, its easy to forget or overlook the decades of research and solutions in information retrieval.
Nonetheless, while embeddings are undoubtedly a powerful tool, they are not the be-all and end-all. First, while they excel at capturing high-level semantic similarity, they may struggle with more specific, keyword-based queries, like when users search for names (e.g., Ilya), acronyms (e.g., RAG), or IDs (e.g., claude-3-sonnet). Keyword-based search, such as BM25, is explicitly designed for this. Finally, after years of keyword-based search, users have likely taken it for granted and may get frustrated if the document they expect to retrieve isnt being returned.
Vector embeddings _do not_ magically solve search. In fact, the heavy lifting is in the step before you re-rank with semantic similarity search. Making a genuine improvement over BM25 or full-text search is hard. — [Aravind Srinivas, CEO Perplexity.ai](https://x.com/AravSrinivas/status/1737886080555446552)
Weve been communicating this to our customers and partners for months now. Nearest Neighbor Search with naive embeddings yields very noisy results and youre likely better off starting with a keyword-based approach. — [Beyang Liu, CTO Sourcegraph](https://twitter.com/beyang/status/1767330006999720318)
Second, its more straightforward to understand why a document was retrieved with keyword search—we can look at the keywords that match the query. In contrast, embedding-based retrieval is less interpretable. Finally, thanks to systems like Lucene and OpenSearch that have been optimized and battle-tested over decades, keyword search is usually more computationally efficient.
In most cases, a hybrid will work best: keyword matching for the obvious matches, and embeddings for synonyms, hypernyms, and spelling errors, as well as multimodality (e.g., images and text). [Shortwave shared how they built their RAG pipeline](https://www.shortwave.com/blog/deep-dive-into-worlds-smartest-email-ai/), including query rewriting, keyword + embedding retrieval, and ranking.
### 1.2.3 Prefer RAG over fine-tuning for new knowledge
Both RAG and fine-tuning can be used to incorporate new information into LLMs and increase performance on specific tasks. However, which should we prioritize?
Recent research suggests RAG may have an edge. [One study](https://arxiv.org/abs/2312.05934) compared RAG against unsupervised finetuning (aka continued pretraining), evaluating both on a subset of MMLU and current events. They found that RAG consistently outperformed fine-tuning for knowledge encountered during training as well as entirely new knowledge. In [another paper](https://arxiv.org/abs/2401.08406), they compared RAG against supervised finetuning on an agricultural dataset. Similarly, the performance boost from RAG was greater than fine-tuning, especially for GPT-4 (see Table 20).
Beyond improved performance, RAG has other practical advantages. First, compared to continuous pretraining or fine-tuning, its easier—and cheaper!—to keep retrieval indices up-to-date. Second, if our retrieval indices have problematic documents that contain toxic or biased content, we can easily drop or modify the offending documents. Consider it an andon cord for [documents that ask us to add glue to pizza](https://x.com/petergyang/status/1793480607198323196).
In addition, the R in RAG provides finer-grained control over how we retrieve documents. For example, if were hosting a RAG system for multiple organizations, by partitioning the retrieval indices, we can ensure that each organization can only retrieve documents from their own index. This ensures that we dont inadvertently expose information from one organization to another.
### 1.2.4 Long-context models wont make RAG obsolete
With Gemini 1.5 providing context windows of up to 10M tokens in size, some have begun to question the future of RAG.
I tend to believe that Gemini 1.5 is significantly overhyped by Sora. A context window of 10M tokens effectively makes most of existing RAG frameworks unnecessary — you simply put whatever your data into the context and talk to the model like usual. Imagine how it does to all the startups / agents / langchain projects where most of the engineering efforts goes to RAG 😅 Or in one sentence: the 10m context kills RAG. Nice work Gemini — [Yao Fu](https://x.com/Francis_YAO_/status/1758935954189115714)
While its true that long contexts will be a game-changer for use cases such as analyzing multiple documents or chatting with PDFs, the rumors of RAGs demise are greatly exaggerated.
First, even with a context size of 10M tokens, wed still need a way to select relevant context. Second, beyond the narrow needle-in-a-haystack eval, weve yet to see convincing data that models can effectively reason over large context sizes. Thus, without good retrieval (and ranking), we risk overwhelming the model with distractors, or may even fill the context window with completely irrelevant information.
Finally, theres cost. During inference, the Transformers time complexity scales linearly with context length. Just because there exists a model that can read your orgs entire Google Drive contents before answering each question doesnt mean thats a good idea. Consider an analogy to how we use RAM: we still read and write from disk, even though there exist compute instances with [RAM running into the tens of terabytes](https://aws.amazon.com/ec2/instance-types/high-memory/).
So dont throw your RAGs in the trash just yet. This pattern will remain useful even as context sizes grow.
## 1.3 Tuning and optimizing workflows
Prompting an LLM is just the beginning. To get the most juice out of them, we need to think beyond a single prompt and embrace workflows. For example, how could we split a single complex task into multiple simpler tasks? When is finetuning or caching helpful with increasing performance and reducing latency/cost? Here, we share proven strategies and real-world examples to help you optimize and build reliable LLM workflows.
### 1.3.1 Step-by-step, multi-turn “flows” can give large boosts
Its common knowledge that decomposing a single big prompt into multiple smaller prompts can achieve better results. For example, [AlphaCodium](https://arxiv.org/abs/2401.08500): By switching from a single prompt to a multi-step workflow, they increased GPT-4 accuracy (pass@5) on CodeContests from 19% to 44%. The workflow includes:
- Reflecting on the problem
- Reasoning on the public tests
- Generating possible solutions
- Ranking possible solutions
- Generating synthetic tests
- Iterating on the solutions on public and synthetic tests.
Small tasks with clear objectives make for the best agent or flow prompts. Its not required that every agent prompt requests structured output, but structured outputs help a lot to interface with whatever system is orchestrating the agents interactions with the environment. Some things to try:
- A tightly-specified, explicit planning step. Also, consider having predefined plans to choose from.
- Rewriting the original user prompts into agent prompts, though this process may be lossy!
- Agent behaviors as linear chains, DAGs, and state machines; different dependency and logic relationships can be more and less appropriate for different scales. Can you squeeze performance optimization out of different task architectures?
- Planning validations; your planning can include instructions on how to evaluate the responses from other agents to make sure the final assembly works well together.
- Prompt engineering with fixed upstream state—make sure your agent prompts are evaluated against a collection of variants of what may have happen before.
### 1.3.2 Prioritize deterministic workflows for now
While AI agents can dynamically react to user requests and the environment, their non-deterministic nature makes them a challenge to deploy. Each step an agent takes has a chance of failing, and the chances of recovering from the error are poor. Thus, the likelihood that an agent completes a multi-step task successfully decreases exponentially as the number of steps increases. As a result, teams building agents find it difficult to deploy reliable agents.
A potential approach is to have agent systems produce deterministic plans which are then executed in a structured, reproducible way. First, given a high-level goal or prompt, the agent generates a plan. Then, the plan is executed deterministically. This allows each step to be more predictable and reliable. Benefits include:
- Generated plans can serve as few-shot samples to prompt or finetune an agent.
- Deterministic execution makes the system more reliable, and thus easier to test and debug. In addition, failures can be traced to the specific steps in the plan.
- Generated plans can be represented as directed acyclic graphs (DAGs) which are easier, relative to a static prompt, to understand and adapt to new situations.
The most successful agent builders may be those with strong experience managing junior engineers because the process of generating plans is similar to how we instruct and manage juniors. We give juniors clear goals and concrete plans, instead of vague open-ended directions, and we should do the same for our agents too.
In the end, the key to reliable, working agents will likely be found in adopting more structured, deterministic approaches, as well as collecting data to refine prompts and finetune models. Without this, well build agents that may work exceptionally well some of the time, but on average, disappoint users.
### 1.3.3 Getting more diverse outputs beyond temperature
Suppose your task requires diversity in an LLMs output. Maybe youre writing an LLM pipeline to suggest products to buy from your catalog given a list of products the user bought previously. When running your prompt multiple times, you might notice that the resulting recommendations are too similar—so you might increase the temperature parameter in your LLM requests.
Briefly, increasing the temperature parameter makes LLM responses more varied. At sampling time, the probability distributions of the next token become flatter, meaning that tokens that are usually less likely get chosen more often. Still, when increasing temperature, you may notice some failure modes related to output diversity. For example, some products from the catalog that could be a good fit may never be output by the LLM. The same handful of products might be overrepresented in outputs, if they are highly likely to follow the prompt based on what the LLM has learned at training time. If the temperature is too high, you may get outputs that reference nonexistent products (or gibberish!)
In other words, increasing temperature does not guarantee that the LLM will sample outputs from the probability distribution you expect (e.g., uniform random). Nonetheless, we have other tricks to increase output diversity. The simplest way is to adjust elements within the prompt. For example, if the prompt template includes a list of items, such as historical purchases, shuffling the order of these items each time theyre inserted into the prompt can make a significant difference.
Additionally, keeping a short list of recent outputs can help prevent redundancy. In our recommended products example, by instructing the LLM to avoid suggesting items from this recent list, or by rejecting and resampling outputs that are similar to recent suggestions, we can further diversify the responses. Another effective strategy is to vary the phrasing used in the prompts. For instance, incorporating phrases like “pick an item that the user would love using regularly” or “select a product that the user would likely recommend to friends” can shift the focus and thereby influence the variety of recommended products.
### 1.3.4 Caching is underrated
Caching saves cost and eliminates generation latency by removing the need to recompute responses for the same input. Furthermore, if a response has previously been guardrailed, we can serve these vetted responses and reduce the risk of serving harmful or inappropriate content.
One straightforward approach to caching is to use unique IDs for the items being processed, such as if were summarizing new articles or [product reviews](https://www.cnbc.com/2023/06/12/amazon-is-using-generative-ai-to-summarize-product-reviews.html). When a request comes in, we can check to see if a summary already exists in the cache. If so, we can return it immediately; if not, we generate, guardrail, and serve it, and then store it in the cache for future requests.
For more open-ended queries, we can borrow techniques from the field of search, which also leverages caching for open-ended inputs. Features like autocomplete, spelling correction, and suggested queries also help normalize user input and thus increase the cache hit rate.
### 1.3.5 When to finetune
We may have some tasks where even the most cleverly designed prompts fall short. For example, even after significant prompt engineering, our system may still be a ways from returning reliable, high-quality output. If so, then it may be necessary to finetune a model for your specific task.
Successful examples include:
- [Honeycombs Natural Language Query Assistant](https://www.honeycomb.io/blog/introducing-query-assistant): Initially, the “programming manual” was provided in the prompt together with n-shot examples for in-context learning. While this worked decently, fine-tuning the model led to better output on the syntax and rules of the domain-specific language.
- [Rechats Lucy](https://www.youtube.com/watch?v=B_DMMlDuJB0): The LLM needed to generate responses in a very specific format that combined structured and unstructured data for the frontend to render correctly. Fine-tuning was essential to get it to work consistently.
Nonetheless, while fine-tuning can be effective, it comes with significant costs. We have to annotate fine-tuning data, finetune and evaluate models, and eventually self-host them. Thus, consider if the higher upfront cost is worth it. If prompting gets you 90% of the way there, then fine-tuning may not be worth the investment. However, if we do decide to finetune, to reduce the cost of collecting human-annotated data, we can [generate and finetune on synthetic data](https://eugeneyan.com/writing/synthetic/), or [bootstrap on open-source data](https://eugeneyan.com/writing/finetuning/).
## 1.4 Evaluation & Monitoring
Evaluating LLMs can be a minefield. The inputs and the outputs of LLMs are arbitrary text, and the tasks we set them to are varied. Nonetheless, rigorous and thoughtful evals are critical—its no coincidence that technical leaders at OpenAI [work on evaluation and give feedback on individual evals](https://twitter.com/eugeneyan/status/1701692908074873036).
Evaluating LLM applications invites a diversity of definitions and reductions: its simply unit testing, or its more like observability, or maybe its just data science. We have found all of these perspectives useful. In the following section, we provide some lessons weve learned about what is important in building evals and monitoring pipelines.
### 1.4.1 Create a few assertion-based unit tests from real input/output samples
Create [unit tests (i.e., assertions)](https://hamel.dev/blog/posts/evals/#level-1-unit-tests) consisting of samples of inputs and outputs from production, with expectations for outputs based on at least three criteria. While three criteria might seem arbitrary, its a practical number to start with; fewer might indicate that your task isnt sufficiently defined or is too open-ended, like a general-purpose chatbot. These unit tests, or assertions, should be triggered by any changes to the pipeline, whether its editing a prompt, adding new context via RAG, or other modifications. This [write-up has an example](https://hamel.dev/blog/posts/evals/#step-1-write-scoped-tests) of an assertion-based test for an actual use case.
Consider beginning with assertions that specify phrases or ideas to either include or exclude in all responses. Also consider checks to ensure that word, item, or sentence counts lie within a range. For other kinds of generation, assertions can look different. [Execution-evaluation](https://www.semanticscholar.org/paper/Execution-Based-Evaluation-for-Open-Domain-Code-Wang-Zhou/1bed34f2c23b97fd18de359cf62cd92b3ba612c3) is a powerful method for evaluating code-generation, wherein you run the generated code and determine that the state of runtime is sufficient for the user-request.
As an example, if the user asks for a new function named foo; then after executing the agents generated code, foo should be callable! One challenge in execution-evaluation is that the agent code frequently leaves the runtime in slightly different form than the target code. It can be effective to “relax” assertions to the absolute most weak assumptions that any viable answer would satisfy.
Finally, using your product as intended for customers (i.e., “dogfooding”) can provide insight into failure modes on real-world data. This approach not only helps identify potential weaknesses, but also provides a useful source of production samples that can be converted into evals.
### 1.4.2 LLM-as-Judge can work (somewhat), but its not a silver bullet
LLM-as-Judge, where we use a strong LLM to evaluate the output of other LLMs, has been met with skepticism by some. (Some of us were initially huge skeptics.) Nonetheless, when implemented well, LLM-as-Judge achieves decent correlation with human judgements, and can at least help build priors about how a new prompt or technique may perform. Specifically, when doing pairwise comparisons (e.g., control vs. treatment), LLM-as-Judge typically gets the direction right though the magnitude of the win/loss may be noisy.
Here are some suggestions to get the most out of LLM-as-Judge: - Use pairwise comparisons: Instead of asking the LLM to score a single output on a [Likert](https://en.wikipedia.org/wiki/Likert_scale) scale, present it with two options and ask it to select the better one. This tends to lead to more stable results. - Control for position bias: The order of options presented can bias the LLMs decision. To mitigate this, do each pairwise comparison twice, swapping the order of pairs each time. Just be sure to attribute wins to the right option after swapping! - Allow for ties: In some cases, both options may be equally good. Thus, allow the LLM to declare a tie so it doesnt have to arbitrarily pick a winner. - Use Chain-of-Thought: Asking the LLM to explain its decision before giving a final preference can increase eval reliability. As a bonus, this allows you to use a weaker but faster LLM and still achieve similar results. Because frequently this part of the pipeline is in batch mode, the extra latency from CoT isnt a problem. - Control for response length: LLMs tend to bias toward longer responses. To mitigate this, ensure response pairs are similar in length.
One particularly powerful application of LLM-as-Judge is checking a new prompting strategy against regression. If you have tracked a collection of production results, sometimes you can rerun those production examples with a new prompting strategy, and use LLM-as-Judge to quickly assess where the new strategy may suffer.
Heres an example of a [simple but effective approach](https://hamel.dev/blog/posts/evals/#automated-evaluation-w-llms) to iterate on LLM-as-Judge, where we simply log the LLM response, judges critique (i.e., CoT), and final outcome. They are then reviewed with stakeholders to identify areas for improvement. Over three iterations, agreement with human and LLM improved from 68% to 94%!
![c model response {"calculations":[{"column":"dur MAX" ","op":"does-not-exist","join_c olumn":" "orders" order" limit" model critque The response is nearly correct, as it is looking for the slowest trace by using and ordering by duration_ms in descending order, which is appropriate for finding the 'slowest' trace. Additionally, filtering with trace.parent_id does-not-exist correctly identifies root spans. However, the query should be grouping by trace.trace_id to ensure that we identify distinct traces, not just the longest individual span. Without the correct grouping, the analysis does not guarantee that the result is a full trace, but merely the longest span. Also, specifying a limit of I is good as it will return the single slowest trace. as reauested. model outcome bad Phillip critique The response is nearly correct, as it is looking for the slowest trace by using MAX(duration_ms) and ordering by duration _ ms in descending order, which is appropriate for finding the 'slowest' trace. Additionally, filtering with trace.parent_id does-not-exist correctly identifies root spans. However, the query should be grouping by trace.trace_id to actually show the slowest trace. Without that grouping, the query only shows the MAX(duration_ms) measurement over time, irrespective of which trace is responsible for that measurement. Phillip outcome bad Phillip revised response {"calculations":[{"column":"dura tion_ms" , "op" "fi Iters" ":" does-not-exist", "join _ column" :""}],"orders":[{"column":"durati scending" Y], "limit" "time_rang agreement TRUE ](Exported%20image%2020240808113930-0.png)
LLM-as-Judge is not a silver bullet though. There are subtle aspects of language where even the strongest models fail to evaluate reliably. In addition, weve found that [conventional classifiers](https://eugeneyan.com/writing/finetuning/) and reward models can achieve higher accuracy than LLM-as-Judge, and with lower cost and latency. For code generation, LLM-as-Judge can be weaker than more direct evaluation strategies like execution-evaluation.
### 1.4.3 The “intern test” for evaluating generations
We like to use the following “intern test” when evaluating generations: If you took the exact input to the language model, including the context, and gave it to an average college student in the relevant major as a task, could they succeed? How long would it take?
If the answer is no because the LLM lacks the required knowledge, consider ways to enrich the context.
If the answer is no and we simply cant improve the context to fix it, then we may have hit a task thats too hard for contemporary LLMs.
If the answer is yes, but it would take a while, we can try to reduce the complexity of the task. Is it decomposable? Are there aspects of the task that can be made more templatized?
If the answer is yes, they would get it quickly, then its time to dig into the data. Whats the model doing wrong? Can we find a pattern of failures? Try asking the model to explain itself before or after it responds, to help you build a theory of mind.
### 1.4.4 Overemphasizing certain evals can hurt overall performance
“When a measure becomes a target, it ceases to be a good measure.” — Goodharts Law.
An example of this is the Needle-in-a-Haystack (NIAH) eval. The original eval helped quantify model recall as context sizes grew, as well as how recall is affected by needle position. However, its been so overemphasized that its featured as [Figure 1 for Gemini 1.5s report](https://arxiv.org/abs/2403.05530). The eval involves inserting a specific phrase (“The special magic {city} number is: {number}”) into a long document which repeats the essays of Paul Graham, and then prompting the model to recall the magic number.
While some models achieve near-perfect recall, its questionable whether NIAH truly reflects the reasoning and recall abilities needed in real-world applications. Consider a more practical scenario: Given the transcript of an hour-long meeting, can the LLM summarize the key decisions and next steps, as well as correctly attribute each item to the relevant person? This task is more realistic, going beyond rote memorization and also considering the ability to parse complex discussions, identify relevant information, and synthesize summaries.
Heres an example of a [practical NIAH eval](https://observablehq.com/@shreyashankar/needle-in-the-real-world-experiments). Using [transcripts of doctor-patient video calls](https://github.com/wyim/aci-bench/tree/main/data/challenge_data), the LLM is queried about the patients medication. It also includes a more challenging NIAH, inserting a phrase for random ingredients for pizza toppings, such as “_The secret ingredients needed to build the perfect pizza are: Espresso-soaked dates, Lemon and Goat cheese._”. Recall was around 80% on the medication task and 30% on the pizza task.
![Avg Recall of Reference Answer in Model Outputs Recall 0.9- 0.8- 0.7- 0.6- 0.5- 0.4- 0.3- 0.2- 0.1- 0.0- dbrx gemini gemma gpt3.5 gpt4 haiku -5 Model mistral opus sonnet ](Exported%20image%2020240808113930-1.png)
Tangentially, an overemphasis on NIAH evals can lead to lower performance on extraction and summarization tasks. Because these LLMs are so finetuned to attend to every sentence, they may start to treat irrelevant details and distractors as important, thus including them in the final output (when they shouldnt!)
This could also apply to other evals and use cases. For example, summarization. An emphasis on factual consistency could lead to summaries that are less specific (and thus less likely to be factually inconsistent) and possibly less relevant. Conversely, an emphasis on writing style and eloquence could lead to more flowery, marketing-type language that could introduce factual inconsistencies.
### 1.4.5 Simplify annotation to binary tasks or pairwise comparisons
Providing open-ended feedback or ratings for model output on a [Likert scale](https://en.wikipedia.org/wiki/Likert_scale) is cognitively demanding. As a result, the data collected is more noisy—due to variability among human raters—and thus less useful. A more effective approach is to simplify the task and reduce the cognitive burden on annotators. Two tasks that work well are binary classifications and pairwise comparisons.
In binary classifications, annotators are asked to make a simple yes-or-no judgment on the models output. They might be asked whether the generated summary is factually consistent with the source document, or whether the proposed response is relevant, or if it contains toxicity. Compared to the Likert scale, binary decisions are more precise, have higher consistency among raters, and lead to higher throughput. This was how [Doordash setup their labeling queues](https://doordash.engineering/2020/08/28/overcome-the-cold-start-problem-in-menu-item-tagging/) for tagging menu items though a tree of yes-no questions.
In pairwise comparisons, the annotator is presented with a pair of model responses and asked which is better. Because its easier for humans to say “A is better than B” than to assign an individual score to either A or B individually, this leads to faster and more reliable annotations (over Likert scales). At a [Llama2 meetup](https://www.youtube.com/watch?v=CzR3OrOkM9w), Thomas Scialom, an author on the Llama2 paper, confirmed that pairwise-comparisons were faster and cheaper than collecting supervised finetuning data such as written responses. The formers cost is $3.5 per unit while the latters cost is $25 per unit.
If youre starting to write labeling guidelines, here are some [reference guidelines](https://eugeneyan.com/writing/labeling-guidelines/) from Google and Bing Search.
### 1.4.6 (Reference-free) evals and guardrails can be used interchangeably
Guardrails help to catch inappropriate or harmful content while evals help to measure the quality and accuracy of the models output. In the case of reference-free evals, they may be considered two sides of the same coin. Reference-free evals are evaluations that dont rely on a “golden” reference, such as a human-written answer, and can assess the quality of output based solely on the input prompt and the models response.
Some examples of these are [summarization evals](https://eugeneyan.com/writing/evals/#summarization-consistency-relevance-length), where we only have to consider the input document to evaluate the summary on factual consistency and relevance. If the summary scores poorly on these metrics, we can choose not to display it to the user, effectively using the eval as a guardrail. Similarly, reference-free [translation evals](https://eugeneyan.com/writing/evals/#translation-statistical--learned-evals-for-quality) can assess the quality of a translation without needing a human-translated reference, again allowing us to use it as a guardrail.
### 1.4.7 LLMs will return output even when they shouldnt
A key challenge when working with LLMs is that theyll often generate output even when they shouldnt. This can lead to harmless but nonsensical responses, or more egregious defects like toxicity or dangerous content. For example, when asked to extract specific attributes or metadata from a document, an LLM may confidently return values even when those values dont actually exist. Alternatively, the model may respond in a language other than English because we provided non-English documents in the context.
While we can try to prompt the LLM to return a “not applicable” or “unknown” response, its not foolproof. Even when the log probabilities are available, theyre a poor indicator of output quality. While log probs indicate the likelihood of a token appearing in the output, they dont necessarily reflect the correctness of the generated text. On the contrary, for instruction-tuned models that are trained to respond to queries and generate coherent response, log probabilities may not be well-calibrated. Thus, while a high log probability may indicate that the output is fluent and coherent, it doesnt mean its accurate or relevant.
While careful prompt engineering can help to some extent, we should complement it with robust guardrails that detect and filter/regenerate undesired output. For example, OpenAI provides a [content moderation API](https://platform.openai.com/docs/guides/moderation) that can identify unsafe responses such as hate speech, self-harm, or sexual output. Similarly, there are numerous packages for [detecting personally identifiable information](https://github.com/topics/pii-detection) (PII). One benefit is that guardrails are largely agnostic of the use case and can thus be applied broadly to all output in a given language. In addition, with precise retrieval, our system can deterministically respond “I dont know” if there are no relevant documents.
A corollary here is that LLMs may fail to produce outputs when they are expected to. This can happen for various reasons, from straightforward issues like long tail latencies from API providers to more complex ones such as outputs being blocked by content moderation filters. As such, its important to consistently log inputs and (potentially a lack of) outputs for debugging and monitoring.
### 1.4.8 Hallucinations are a stubborn problem
Unlike content safety or PII defects which have a lot of attention and thus seldom occur, factual inconsistencies are stubbornly persistent and more challenging to detect. Theyre more common and occur at a baseline rate of 5 - 10%, and from what weve learned from LLM providers, it can be challenging to get it below 2%, even on simple tasks such as summarization.
To address this, we can combine prompt engineering (upstream of generation) and factual inconsistency guardrails (downstream of generation). For prompt engineering, techniques like CoT help reduce hallucination by getting the LLM to explain its reasoning before finally returning the output. Then, we can apply a [factual inconsistency guardrail](https://eugeneyan.com/writing/finetuning/) to assess the factuality of summaries and filter or regenerate hallucinations. In some cases, hallucinations can be deterministically detected. When using resources from RAG retrieval, if the output is structured and identifies what the resources are, you should be able to manually verify theyre sourced from the input context.
# 2 Operational: Day-to-day and org concerns
## 2.1 Data
Just as the quality of ingredients determines the dishs taste, the quality of input data constrains the performance of machine learning systems. In addition, output data is the only way to tell whether the product is working or not. All the authors focus tightly on the data, looking at inputs and outputs for several hours a week to better understand the data distribution: its modes, its edge cases, and the limitations of models of it.
### 2.1.1 Check for development-prod skew
A common source of errors in traditional machine learning pipelines is _train-serve skew_. This happens when the data used in training differs from what the model encounters in production. Although we can use LLMs without training or fine-tuning, hence theres no training set, a similar issue arises with development-prod data skew. Essentially, the data we test our systems on during development should mirror what the systems will face in production. If not, we might find our production accuracy suffering.
LLM development-prod skew can be categorized into two types: structural and content-based. Structural skew includes issues like formatting discrepancies, such as differences between a JSON dictionary with a list-type value and a JSON list, inconsistent casing, and errors like typos or sentence fragments. These errors can lead to unpredictable model performance because different LLMs are trained on specific data formats, and prompts can be highly sensitive to minor changes. Content-based or “semantic” skew refers to differences in the meaning or context of the data. 
As in traditional ML, its useful to periodically measure skew between the LLM input/output pairs. Simple metrics like the length of inputs and outputs or specific formatting requirements (e.g., JSON or XML) are straightforward ways to track changes. For more “advanced” drift detection, consider clustering embeddings of input/output pairs to detect semantic drift, such as shifts in the topics users are discussing, which could indicate they are exploring areas the model hasnt been exposed to before. 
When testing changes, such as prompt engineering, ensure that hold-out datasets are current and reflect the most recent types of user interactions. For example, if typos are common in production inputs, they should also be present in the hold-out data. Beyond just numerical skew measurements, its beneficial to perform qualitative assessments on outputs. Regularly reviewing your models outputs—a practice colloquially known as “vibe checks”—ensures that the results align with expectations and remain relevant to user needs. Finally, incorporating nondeterminism into skew checks is also useful—by running the pipeline multiple times for each input in our testing dataset and analyzing all outputs, we increase the likelihood of catching anomalies that might occur only occasionally.
### 2.1.2 Look at samples of LLM inputs and outputs every day
LLMs are dynamic and constantly evolving. Despite their impressive zero-shot capabilities and often delightful outputs, their failure modes can be highly unpredictable. For custom tasks, regularly reviewing data samples is essential to developing an intuitive understanding of how LLMs perform.
Input-output pairs from production are the “real things, real places” (_genchi genbutsu_) of LLM applications, and they cannot be substituted. [Recent research](https://arxiv.org/abs/2404.12272) highlighted that developers perceptions of what constitutes “good” and “bad” outputs shift as they interact with more data (i.e., _criteria drift_). While developers can come up with some criteria upfront for evaluating LLM outputs, these predefined criteria are often incomplete. For instance, during the course of development, we might update the prompt to increase the probability of good responses and decrease the probability of bad ones. This iterative process of evaluation, reevaluation, and criteria update is necessary, as its difficult to predict either LLM behavior or human preference without directly observing the outputs.
To manage this effectively, we should log LLM inputs and outputs. By examining a sample of these logs daily, we can quickly identify and adapt to new patterns or failure modes. When we spot a new issue, we can immediately write an assertion or eval around it. Similarly, any updates to failure mode definitions should be reflected in the evaluation criteria. These “vibe checks” are signals of bad outputs; code and assertions operationalize them. Finally, this attitude must be socialized, for example by adding review or annotation of inputs and outputs to your on-call rotation.
## 2.2 Working with models
With LLM APIs, we can rely on intelligence from a handful of providers. While this is a boon, these dependencies also involve trade-offs on performance, latency, throughput, and cost. Also, as newer, better models drop (almost every month in the past year), we should be prepared to update our products as we deprecate old models and migrate to newer models. In this section, we share our lessons from working with technologies we dont have full control over, where the models cant be self-hosted and managed.
### 2.2.1 Generate structured output to ease downstream integration
For most real-world use cases, the output of an LLM will be consumed by a downstream application via some machine-readable format. For example, [ReChat](https://www.youtube.com/watch?v=B_DMMlDuJB0), a real-estate CRM, required structured responses for the front end to render widgets. Similarly, [Boba](https://martinfowler.com/articles/building-boba.html), a tool for generating product strategy ideas, needed structured output with fields for title, summary, plausibility score, and time horizon. Finally, LinkedIn shared about [constraining the LLM to generate YAML](https://www.linkedin.com/blog/engineering/generative-ai/musings-on-building-a-generative-ai-product), which is then used to decide which skill to use, as well as provide the parameters to invoke the skill.
This application pattern is an extreme version of Postels Law: be liberal in what you accept (arbitrary natural language) and conservative in what you send (typed, machine-readable objects). As such, we expect it to be extremely durable.
Currently, [Instructor](https://github.com/jxnl/instructor) and [Outlines](https://github.com/outlines-dev/outlines) are the de facto standards for coaxing structured output from LLMs. If youre using an LLM API (e.g., Anthropic, OpenAI), use Instructor; if youre working with a self-hosted model (e.g., Huggingface), use Outlines.
### 2.2.2 Migrating prompts across models is a pain in the ass
Sometimes, our carefully crafted prompts work superbly with one model but fall flat with another. This can happen when were switching between various model providers, as well as when we upgrade across versions of the same model. 
For example, Voiceflow found that [migrating from gpt-3.5-turbo-0301 to gpt-3.5-turbo-1106 led to a 10% drop](https://www.voiceflow.com/blog/how-much-do-chatgpt-versions-affect-real-world-performance) on their intent classification task. (Thankfully, they had evals!) Similarly, [GoDaddy observed a trend in the positive direction](https://www.godaddy.com/resources/news/llm-from-the-trenches-10-lessons-learned-operationalizing-models-at-godaddy#h-3-prompts-aren-t-portable-across-models), where upgrading to version 1106 narrowed the performance gap between gpt-3.5-turbo and gpt-4. (Or, if youre a glass-half-full person, you might be disappointed that gpt-4s lead was reduced with the new upgrade)
Thus, if we have to migrate prompts across models, expect it to take more time than simply swapping the API endpoint. Dont assume that plugging in the same prompt will lead to similar or better results. Also, having reliable, automated evals helps with measuring task performance before and after migration, and reduces the effort needed for manual verification.
### 2.2.3 Version and pin your models
In any machine learning pipeline, “[changing anything changes everything](https://papers.nips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html)”. This is particularly relevant as we rely on components like large language models (LLMs) that we dont train ourselves and that can change without our knowledge.
Fortunately, many model providers offer the option to “pin” specific model versions (e.g., gpt-4-turbo-1106). This enables us to use a specific version of the model weights, ensuring they remain unchanged. Pinning model versions in production can help avoid unexpected changes in model behavior, which could lead to customer complaints about issues that may crop up when a model is swapped, such as overly verbose outputs or other unforeseen failure modes.
Additionally, consider maintaining a shadow pipeline that mirrors your production setup but uses the latest model versions. This enables safe experimentation and testing with new releases. Once youve validated the stability and quality of the outputs from these newer models, you can confidently update the model versions in your production environment.
### 2.2.4 Choose the smallest model that gets the job done
When working on a new application, its tempting to use the biggest, most powerful model available. But once weve established that the task is technically feasible, its worth experimenting if a smaller model can achieve comparable results.
The benefits of a smaller model are lower latency and cost. While it may be weaker, techniques like chain-of-thought, n-shot prompts, and in-context learning can help smaller models punch above their weight. Beyond LLM APIs, fine-tuning our specific tasks can also help increase performance.
Taken together, a carefully crafted workflow using a smaller model can often match, or even surpass, the output quality of a single large model, while being faster and cheaper. For example, this [tweet](https://twitter.com/mattshumer_/status/1770823530394833242) shares anecdata of how Haiku + 10-shot prompt outperforms zero-shot Opus and GPT-4. In the long term, we expect to see more examples of [flow-engineering](https://twitter.com/karpathy/status/1748043513156272416) with smaller models as the optimal balance of output quality, latency, and cost.
As another example, take the humble classification task. Lightweight models like DistilBERT (67M parameters) are a surprisingly strong baseline. The 400M parameter DistilBART is another great option—when finetuned on open-source data, it could [identify hallucinations with an ROC-AUC of 0.84](https://eugeneyan.com/writing/finetuning/), surpassing most LLMs at less than 5% of latency and cost.
The point is, dont overlook smaller models. While its easy to throw a massive model at every problem, with some creativity and experimentation, we can often find a more efficient solution. 
## 2.3 Product
While new technology offers new possibilities, the principles of building great products are timeless. Thus, even if were solving new problems for the first time, we dont have to reinvent the wheel on product design. Theres a lot to gain from grounding our LLM application development in solid product fundamentals, allowing us to deliver real value to the people we serve.
### 2.3.1 Involve design early and often
Having a designer will push you to understand and think deeply about how your product can be built and presented to users. We sometimes stereotype designers as folks who take things and make them pretty. But beyond just the user interface, they also rethink how the user experience can be improved, even if it means breaking existing rules and paradigms.
Designers are especially gifted at reframing the users needs into various forms. Some of these forms are more tractable to solve than others, and thus, they may offer more or fewer opportunities for AI solutions. Like many other products, building AI products should be centered around the job to be done, not the technology that powers them.
Focus on asking yourself: “What job is the user asking this product to do for them? Is that job something a chatbot would be good at? How about autocomplete? Maybe something different!” Consider the existing [design patterns](https://www.tidepool.so/blog/emerging-ux-patterns-for-generative-ai-apps-copilots) and how they relate to the job-to-be-done. These are the invaluable assets that designers add to your teams capabilities.
### 2.3.2 Design your UX for Human-In-The-Loop
One way to get quality annotations is to integrate Human-in-the-Loop (HITL) into the user experience (UX). By allowing users to provide feedback and corrections easily, we can improve the immediate output and collect valuable data to improve our models.
Imagine an e-commerce platform where users upload and categorize their products. There are several ways we could design the UX:
- The user manually selects the right product category; an LLM periodically checks new products and corrects miscategorization on the backend.
- The user doesnt select any category at all; an LLM periodically categorizes products on the backend (with potential errors).
- An LLM suggests a product category in real-time, which the user can validate and update as needed.
While all three approaches involve an LLM, they provide very different UXes. The first approach puts the initial burden on the user and has the LLM acting as a post-processing check. The second requires zero effort from the user but provides no transparency or control. The third strikes the right balance. By having the LLM suggest categories upfront, we reduce cognitive load on the user and they dont have to learn our taxonomy to categorize their product! At the same time, by allowing the user to review and edit the suggestion, they have the final say in how their product is classified, putting control firmly in their hands. As a bonus, the third approach creates a [natural feedback loop for model improvement](https://eugeneyan.com/writing/llm-patterns/#collect-user-feedback-to-build-our-data-flywheel). Suggestions that are good are accepted (positive labels) and those that are bad are updated (negative followed by positive labels).
This pattern of suggestion, user validation, and data collection is commonly seen in several applications:
- Coding assistants: Where users can accept a suggestion (strong positive), accept and tweak a suggestion (positive), or ignore a suggestion (negative)
- Midjourney: Where users can choose to upscale and download the image (strong positive), vary an image (positive), or generate a new set of images (negative)
- Chatbots: Where users can provide thumbs up (positive) or thumbs down (negative) on responses, or choose to regenerate a response if it was really bad (strong negative).
Feedback can be explicit or implicit. Explicit feedback is information users provide in response to a request by our product; implicit feedback is information we learn from user interactions without needing users to deliberately provide feedback. Coding assistants and Midjourney are examples of implicit feedback while thumbs up and thumb downs are explicit feedback. If we design our UX well, like coding assistants and Midjourney, we can collect plenty of implicit feedback to improve our product and models.
### 2.3.3 Prioritize your hierarchy of needs ruthlessly
As we think about putting our demo into production, well have to think about the requirements for:
- Reliability: 99.9% uptime, adherence to structured output
- Harmlessness: Not generate offensive, NSFW, or otherwise harmful content
- Factual consistency: Being faithful to the context provided, not making things up
- Usefulness: Relevant to the users needs and request
- Scalability: Latency SLAs, supported throughput
- Cost: Because we dont have unlimited budget
- And more: Security, privacy, fairness, GDPR, DMA, etc, etc.
If we try to tackle all these requirements at once, were never going to ship anything. Thus, we need to prioritize. Ruthlessly. This means being clear what is non-negotiable (e.g., reliability, harmlessness) without which our product cant function or wont be viable. Its all about identifying the minimum lovable product. We have to accept that the first version wont be perfect, and just launch and iterate.
### 2.3.4 Calibrate your risk tolerance based on the use case
When deciding on the language model and level of scrutiny of an application, consider the use case and audience. For a customer-facing chatbot offering medical or financial advice, well need a very high bar for safety and accuracy. Mistakes or bad output could cause real harm and erode trust. But for less critical applications, such as a recommender system, or internal-facing applications like content classification or summarization, excessively strict requirements only slow progress without adding much value.
This aligns with a recent [a16z report](https://a16z.com/generative-ai-enterprise-2024/) showing that many companies are moving faster with internal LLM applications compared to external ones. By experimenting with AI for internal productivity, organizations can start capturing value while learning how to manage risk in a more controlled environment. Then, as they gain confidence, they can expand to customer-facing use cases.
## 2.4 Team & Roles
No job function is easy to define, but writing a job description for the work in this new space is more challenging than others. Well forgo venn diagrams of intersecting job titles, or suggestions for job descriptions. We will, however, submit to the existence of a new role—the AI engineer—and discuss its place. Importantly, well discuss the rest of the team and how responsibilities should be assigned.
### 2.4.1 Focus on process, not tools
When faced with new paradigms, such as LLMs, software engineers tend to favor tools. As a result, we overlook the problem and process the tool was supposed to solve. In doing so, many engineers assume accidental complexity, which has negative consequences for the teams long-term productivity.
For example, [this write-up](https://hamel.dev/blog/posts/prompt/) discusses how certain tools can automatically create prompts for large language models. It argues (rightfully IMHO) that engineers who use these tools without first understanding the problem-solving methodology or process end up taking on unnecessary technical debt.
In addition to accidental complexity, tools are often underspecified. For example, there is a growing industry of LLM evaluation tools that offer “LLM Evaluation In A Box” with generic evaluators for toxicity, conciseness, tone, etc. We have seen many teams adopt these tools without thinking critically about the specific failure modes of their domains. Contrast this to EvalGen. It focuses on teaching users the process of creating domain-specific evals by deeply involving the user each step of the way, from specifying criteria, to labeling data, to checking evals. The software leads the user through a workflow that looks like this:
![Prompt Node Infer You be doing named entity recognition (NER). Extract up to 3 well-known entities fron the following tweet: {tweet_fu For each entity, write one sentence describing the person or entity. ALI the entities you extract should be found in a knowledge base Like Wikipedia, SO don't make up Nurn responses per prompt: Multi -Evaluator criteria from my context you extract should' Let me specify criteria manually Let an Al help you generate criteria and implement evaluation functions. Is this or ? - Bravotv: A television network that focuses on reality TV shows, including popular f tike The Real Housewives. - BravoWWHL: Stands for Bravo's Watch What Happens Live, a late—night talk hosted by Andy Cohen that features celebrity games, and discussions about Bravo's reality TV shows. — Paris Hilton: A well—known socialite, businesswoman, and media personality, known for her appearance on the reality TV show The Simple Life and her work as a Singer, actress, and entrepreneur. ntities I'm tired Type a criteria to add, then press Enter: Markdown Format The response should be in Markdown format. NO Made up Entities There shouldn't be any made up entities in the response, Grade some responses first Grade Some responses first. to help sn•urselt identify criteria. The Al will incorporate your grades in its criteria suggestions. Suggest more Bulleted List The response should contain a bulleted list. NO Hashtags The response should not extract hashtags as entities. "_text STOLE WY HOUSE *justdoit #juststealit uravOtv @BravON4L https://t.co.'8xShKrøSYq Prompt You be doing nard entity recognition (NER). Extract up to 3 entities f rn the tueet : you STOLE MY GOD""' HOUSE *justdolt •juststeatit aparisHilton https://t.co'8xShKreSYq For each entity, write one sentence describing the person or entity. Alt the I'm done. Implement it! Coverage ot Bad Reswtses 77.78% Ento Stuw False Failure Rate 28.57% ](Exported%20image%2020240808113930-2.png)
[Shankar, S., et al. (2024). Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences. Retrieved from https://arxiv.org/abs/2404.12272](https://arxiv.org/abs/2404.12272)
EvalGen guides the user through a best practice of crafting LLM evaluations, namely:
1. Defining domain-specific tests (bootstrapped automatically from the prompt). These are defined as either assertions with code or with LLM-as-a-Judge.
2. The importance of aligning the tests with human judgment, so that the user can check that the tests capture the specified criteria.
3. Iterating on your tests as the system (prompts, etc) changes. 
EvalGen provides developers with a mental model of the evaluation building process without anchoring them to a specific tool. We have found that after providing AI Engineers with this context, they often decide to select leaner tools or build their own.  
There are too many components of LLMs beyond prompt writing and evaluations to list exhaustively here.  However, it is important that AI Engineers seek to understand the processes before adopting tools.
### 2.4.2 Always be experimenting
ML products are deeply intertwined with experimentation. Not only the A/B, Randomized Control Trials kind, but the frequent attempts at modifying the smallest possible components of your system, and doing offline evaluation. The reason why everyone is so hot for evals is not actually about trustworthiness and confidence—its about enabling experiments! The better your evals, the faster you can iterate on experiments, and thus the faster you can converge on the best version of your system. 
Its common to try different approaches to solving the same problem because experimentation is so cheap now. The high-cost of collecting data and training a model is minimized—prompt engineering costs little more than human time. Position your team so that everyone is taught the basics of prompt engineering. This encourages everyone to experiment and leads to diverse ideas from across the organization.
Additionally, dont only experiment to explore—also use them to exploit! Have a working version of a new task? Consider having someone else on the team approach it differently. Try doing it another way thatll be faster. Investigate prompt techniques like Chain-of-Thought or Few-Shot to make it higher quality. Dont let your tooling hold you back on experimentation; if it is, rebuild it, or buy something to make it better. 
Finally, during product/project planning, set aside time for building evals and running multiple experiments. Think of the product spec for engineering products, but add to it clear criteria for evals. And during roadmapping, dont underestimate the time required for experimentation—expect to do multiple iterations of development and evals before getting the green light for production.
### 2.4.3 Empower everyone to use new AI technology
As generative AI increases in adoption, we want the entire team—not just the experts—to understand and feel empowered to use this new technology. Theres no better way to develop intuition for how LLMs work (e.g., latencies, failure modes, UX) than to, well, use them. LLMs are relatively accessible: You dont need to know how to code to improve performance for a pipeline, and everyone can start contributing via prompt engineering and evals.
A big part of this is education. It can start as simple as the basics of prompt engineering, where techniques like n-shot prompting and CoT help condition the model towards the desired output. Folks who have the knowledge can also educate about the more technical aspects, such as how LLMs are autoregressive in nature. In other words, while input tokens are processed in parallel, output tokens are generated sequentially. As a result, latency is more a function of output length than input length—this is a key consideration when designing UXes and setting performance expectations.
We can also go further and provide opportunities for hands-on experimentation and exploration. A hackathon perhaps? While it may seem expensive to have an entire team spend a few days hacking on speculative projects, the outcomes may surprise you. We know of a team that, through a hackathon, accelerated and almost completed their three-year roadmap within a year. Another team had a hackathon that led to paradigm shifting UXes that are now possible thanks to LLMs, which are now prioritized for the year and beyond.
### 2.4.4 Dont fall into the trap of “AI Engineering is all I need”
As new job titles are coined, there is an initial tendency to overstate the capabilities associated with these roles. This often results in a painful correction as the actual scope of these jobs becomes clear. Newcomers to the field, as well as hiring managers, might make exaggerated claims or have inflated expectations. Notable examples over the last decade include:
- Data Scientist: “[someone who is better at statistics than any software engineer and better at software engineering than any statistician](https://x.com/josh_wills/status/198093512149958656).”  
- Machine Learning Engineer (MLE): a software engineering-centric view of machine learning 
Initially, many assumed that data scientists alone were sufficient for data-driven projects. However, it became apparent that data scientists must collaborate with software and data engineers to develop and deploy data products effectively. 
This misunderstanding has shown up again with the new role of AI Engineer, with some teams believing that AI Engineers are all you need. In reality, building machine learning or AI products requires a [broad array of specialized roles](https://papers.nips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html). Weve consulted with more than a dozen companies on AI products and have consistently observed that they fall into the trap of believing that “AI Engineering is all you need.” As a result, products often struggle to scale beyond a demo as companies overlook crucial aspects involved in building a product.
For example, evaluation and measurement are crucial for scaling a product beyond vibe checks. The skills for effective evaluation align with some of the strengths traditionally seen in machine learning engineers—a team composed solely of AI Engineers will likely lack these skills. Co-author Hamel Husain illustrates the importance of these skills in his recent work around detecting [data drift](https://github.com/hamelsmu/ft-drift) and [designing domain-specific evals](https://hamel.dev/blog/posts/evals/).
Here is a rough progression of the types of roles you need, and when youll need them, throughout the journey of building an AI product:
1. First, focus on building a product. This might include an AI engineer, but it doesnt have to. AI Engineers are valuable for prototyping and iterating quickly on the product (UX, plumbing, etc). 
2. Next, create the right foundations by instrumenting your system and collecting data. Depending on the type and scale of data, you might need platform and/or data engineers. You must also have systems for querying and analyzing this data to debug issues.
3. Next, you will eventually want to optimize your AI system. This doesnt necessarily involve training models. The basics include steps like designing metrics, building evaluation systems, running experiments, optimizing RAG retrieval, debugging stochastic systems, and more. MLEs are really good at this (though AI engineers can pick them up too). It usually doesnt make sense to hire an MLE unless you have completed the prerequisite steps.
Aside from this, you need a domain expert at all times. At small companies, this would ideally be the founding team—and at bigger companies, product managers can play this role. Being aware of the progression and timing of roles is critical. Hiring folks at the wrong time (e.g., [hiring an MLE too early](https://jxnl.co/writing/2024/04/08/hiring-mle-at-early-stage-companies/)) or building in the wrong order is a waste of time and money, and causes churn.  Furthermore, regularly checking in with an MLE (but not hiring them full-time) during phases 1-2 will help the company build the right foundations. 
# 3 Strategic: Long-term business strategy (pending)
PENDING RELEASE (tentatively 6th June)
# 4 Stay In Touch
If you found this useful and want updates on write-ups, courses, and activities, subscribe below.
You can also find our individual contact information on our [about page](https://applied-llms.org/about.html).
## 4.1 Acknowledgements
This series started as a conversation in a group chat, where Bryan quipped that he was inspired to write “A Year of AI Engineering”. Then, ✨magic✨ happened, and we were all inspired to chip in and share what weve learned so far.
The authors would like to thank Eugene for leading the bulk of the document integration and overall structure in addition to a large proportion of the lessons. Additionally, for primary editing responsibilities and document direction. The authors would like to thank Bryan for the spark that led to this writeup, restructuring the write-up into tactical, operational, and strategic sections and their intros, and for pushing us to think bigger on how we could reach and help the community. The authors would like to thank Charles for his deep dives on cost and LLMOps, as well as weaving the lessons to make them more coherent and tighter—you have him to thank for this being 30 instead of 40 pages! The authors thank Hamel and Jason for their insights from advising clients and being on the front lines, for their broad generalizable learnings from clients, and for deep knowledge of tools. And finally, thank you Shreya for reminding us of the importance of evals and rigorous production practices and for bringing her research and original results.
Finally, we would like to thank all the teams who so generously shared your challenges and lessons in your own write-ups which weve referenced throughout this series, along with the AI communities for your vibrant participation and engagement with this group.
## 4.2 About the authors
See the [about page](https://applied-llms.org/about.html) for more information on the authors.
If you found this useful, please cite this write-up as:
Yan, Eugene, Bryan Bischof, Charles Frye, Hamel Husain, Jason Liu, and Shreya Shankar. 2024. Applied LLMs - What Weve Learned From A Year of Building with LLMs. Applied LLMs. 8 June 2024. [https://applied-llms.org/](https://applied-llms.org/).
or
@article{AppliedLLMs2024, title = {What We've Learned From A Year of Building with LLMs}, author = {Yan, Eugene and Bischof, Bryan and Frye, Charles and Husain, Hamel and Liu, Jason and Shankar, Shreya}, journal = {Applied LLMs}, year = {2024}, month = {Jun}, url = {https://applied-llms.org/}}

View File

@@ -0,0 +1,952 @@
Clipped from: [https://codeblog.jonskeet.uk/category/edulinq/](https://codeblog.jonskeet.uk/category/edulinq/)
# Category Archives: Edulinq
[Books](https://codeblog.jonskeet.uk/category/books/), [C#](https://codeblog.jonskeet.uk/category/csharp/), [Edulinq](https://codeblog.jonskeet.uk/category/edulinq/), [LINQ](https://codeblog.jonskeet.uk/category/linq/)
# [Edulinq the e-book](https://codeblog.jonskeet.uk/2011/03/18/edulinq-the-e-book/)
[March 18, 2011](https://codeblog.jonskeet.uk/2011/03/18/edulinq-the-e-book/) [jonskeet](https://codeblog.jonskeet.uk/author/jonskeet/) [15 Comments](https://codeblog.jonskeet.uk/2011/03/18/edulinq-the-e-book/#comments)
Im pleased to announce that Ive made a first pass at converting the blog posts in the Edulinq series into e-books.
Im using [Calibre](http://calibre-ebook.com/) to convert to PDF and e-book format. I still have a way to go, but theyre at least readable. The Kindle version (MOBI format) is working somewhat better than the PDF version at the moment, which surprises me. In particular, although hyperlinks are _displaying_ in the PDF, they dont seem to be working whereas at least the internal links in the Kindle format are working.
Ill no doubt try to improve things over time, but Ive put these early attempts up on the [root page](https://github.com/jskeet/edulinq) of the Edulinq project site. Ive also put all the blog posts up as HTML in the sites source control; that means you can [browse the latest version](https://github.com/jskeet/edulinq/tree/master/posts) directly. It also means if you sync the source control, download Calibre yourself and add “index.html” as a new e-book in the Calibre library, you can play with the conversion yourself and help me improve things. Feedback about problems is welcome; feedback including the fix is even better :)
[C#](https://codeblog.jonskeet.uk/category/csharp/), [Edulinq](https://codeblog.jonskeet.uk/category/edulinq/), [LINQ](https://codeblog.jonskeet.uk/category/linq/)
# [Reimplementing LINQ to Objects: Part 44 Aspects of Design](https://codeblog.jonskeet.uk/2011/02/21/reimplementing-linq-to-objects-part-44-aspects-of-design/)
[February 21, 2011](https://codeblog.jonskeet.uk/2011/02/21/reimplementing-linq-to-objects-part-44-aspects-of-design/) [jonskeet](https://codeblog.jonskeet.uk/author/jonskeet/) [3 Comments](https://codeblog.jonskeet.uk/2011/02/21/reimplementing-linq-to-objects-part-44-aspects-of-design/#comments)
I promised a post on some questions of design that are raised by LINQ to Objects. I suspect that most of these have already been covered in other posts, but it may well be helpful to talk about them here too. This time Ive thought about it particularly from the point of view of how other APIs can be built on some of the same design principles, and the awkward choices that LINQ has thrown up.
### The power of composability and immutability
Perhaps the most important aspect of LINQ which Id love other API designers to take on board is that of how complicated queries are constructed from lots of little building blocks. What makes it particularly elegant is that the result of applying each building block is unchanged by anything else you do to it afterwards.
LINQ doesnt _enforce_ immutability of course you can start off with a mutable list and change its content at any time, for example, or change the properties of one of the objects referenced within it, or pass in a delegate with side-effects but _LINQ itself_ wont introduce side-effects.
The Task-based Asynchronous Pattern takes a similar approach, allowing composable building blocks of tasks. Ive seen this pattern in various guises over the years if you find yourself thinking in terms of a pipeline of some kind, it may well be appropriate, especially if each state in the pipeline emits the same type as it consumes.
General immutability is a somewhat different design trait of course, but one which can make _such_ a difference. The java.util.{Date,Calendar} classes are horrible, not least because theyre mutable you can never stash a value away without being concerned that it may get changed by something else. [Joda Time](http://joda-time.sf.net/) has _some_ mutable implementations, but typically the immutable classes are used in a fluent way. Of course, .NET uses value types for various core types to start with, but also makes TimeZoneInfo immutable. For genuine "values" I would highly encourage API designers to at least _strongly consider_ immutable types. Theyre not _always_ appropriate by any means, but they can be hugely useful where they fit nicely.
### Extension methods on interfaces
Its no surprise that extension methods are heavily used in LINQ, given that they were effectively introduced into the language in order to enable LINQ in the first place. However, they do work particularly well with interfaces as a way of adding common behaviour.
It also plays very nicely with the pipeline pattern above for creating pipelines in a fluent manner. Even if you just create extension methods which call a constructor to wrap/compose the previous stage in the pipeline, you can still end up with more readable code.
One _problem_ with this is that you cant "override" behaviour in particular implementations or interfaces which extend the original one which is why Enumerable.ElementAt() has to detect that a sequence is actually a list, for example. If interfaces allowed method implementations, this wouldnt be as much of a problem in the situation where youre in control of the interface I wouldnt be at all surprised to find that as a feature of C#s successor.
The lack of extension properties is also a bit of a handicap in some places, although not as many as one might expect at first glance. For example, even if we _could_ have made Enumerable.Count() a property, would it have been a good idea to do so? Properties give a natural expectation of speed, and Count() is usually an O(n) operation.
### Delegates for custom behaviour
In .NET 1.0 and 1.1, most developers used delegates for two purposes:
- Handling events in UIs
- Passing around behaviour to be executed in a different thread (either via Control.Invoke, or new Thread(ThreadStart), or ThreadPool.QueueUserWorkItem).
.NET 2.0 increased the range of uses of delegates somewhat, particularly with List.ConvertAll and the ability to create delegates relatively easily using anonymous methods.
However, LINQ _really_ brought them into the mainstream. If youre building an API which benefits from _small_ pieces of custom behaviour, delegates can be a real boon. More complicated behaviour is still often best represented via an interface, and _sometimes_ its worth having both interface and delegate representations, like Comparison<T> and IComparable<T>. Its generally easy to convert between the two especially if you use a method group conversion from an interface implementations method to the delegate type.
### Laziness
One aspect of LINQ which is both a blessing and a curse is its laziness, both in terms of deferred execution (not reading from the input sequence _at all_ until the result sequence is read) and in terms of streaming the data (only reading as much information from the input sequence as is required to answer the immediate needs of the caller).
This is great in various ways, particularly as it means you can build a complex query and use it multiple times, sometimes as a basis for other queries, knowing that it wont actually do anything until you ask for real results. It also means that you can iterate over huge data sets, so long as youre careful.
On the other hand, it leads to subtle issues over when code is actually executing, makes debugging harder to understand, makes it easier to accidentally change the values of captured variables between the point at which you create the query and the point at which you execute it, and basically messes with your head. This is probably the aspect of LINQ which confuses newbies more than any other.
Im not saying it was the wrong decision for LINQ but I _would_ caution API designers to think carefully before introducing laziness, and to document it _really_ thoroughly. Likewise if your API might end up returning a result which is "gradually evaluated" (streaming data etc), this should be made clear.
### When names collide: options for consistency
Just in case youve forgotten, this is irritation with the meaning of source.Contains(element). In order to check whether a sequence contains an element or not, there has to be some idea of equality for example, if youre trying to find one string in a sequence of strings, are you trying to match in a case sensitive manner or not?
Theres an overload for Enumerable.Contains which allows you to specify the equality comparer to use, but the question is what should happen when you let the implementation pick the comparer.
For _every_ other method in Enumerable, the default equality comparer for the sequence type i.e. EqualityComparer<TSource>.Default is picked. That sounds like source.Contains(element) should use the element types default comparer too, right? Well, in some cases thats what will happen… but not if the source implements ICollection<T>, which has its own Contains method which doesnt take an equality comparer. If thats the case, LINQ to Objects delegates to the collections Contains method.
So, we have three kinds of consistency here:
- Consistency of compile-time type: it would be nice if the behaviour of source.Contains(element) was the same whether "source" is of type IEnumerable<T> and ICollection<T>
- Consistency of API: it would be nice if Contains behaved the same way as other methods which have overloads with and without equality comparers
- Consistency of model: if you consider "source" to be just a sequence of elements, it shouldnt affect the _result_ (not just the speed) if the object actually implements ICollection<T>
I should point out that this will _only_ be a problem if the collection uses a different notion of equality to the default equality comparer for the type. The most obvious example of this is if you have a HashSet<string> which uses a case-insensitive equality comparer. But its still a valid concern.
So what should the API designer do in this kind of case? Admittedly LINQ to Objects is already in a slightly unusual position as its based on an existing interface with known and _very common_ interfaces extending that core one… its less likely to come up with other APIs. However, I think it _might_ be enough of a smell to suggest that changing the name of the method to "ContainsElement" or something similar would be worthwhile. Its unfortunate that "Contains" really _is_ the obvious choice…
This issue raises another aspect of API decision Ive considered in the past… if theres a common way of doing something in the framework youre building on top of, but you consider it to be broken, should you abide by that breakage for the sake of familiarity and consistency, or should you strive to be as "clean" as possible? I think it needs to be considered on a case-by-case basis, but I _suspect_ I would usually come down on the side of cleanliness.
### Documentation details
Almost all APIs are badly documented its a fact of life, even with some of the best APIs Ive worked with. I doubt that Noda Time will be a shining example either. However, at the risk of being hypocritical Ill say that documentation is worth spending significant thought on. Not just the time taken to document your code but the time taken to consider what you want to guarantee, what should be left unstated, and what should be _explicitly_ left open.
For example, theres no indication in the documentation of Cast that it will _sometimes_ return the original source value, nor in its companion OfType method that that will _never_ return the original source reference. This might be important to someone why not state it? Its possible to state the _possibility_ without saying what cases it applies to of course, leaving some wiggle room in the future. You might consider some of the optimizations in the same way when should an optimization be documented and when should it be implicit? Sometimes it can make a difference beyond just performance, even if only in "odd" situations (such as a predicate throwing an exception).
If youre used to defensive coding with Code Contracts, its much the same type of decision and again, its similar to deciding whether a method should return IEnumerable<T>, IList<T> or List<T>. Theres a balance between caller convenience, design cleanliness (where you only want to _emphasize_ one interface aspect, even if it also _happens_ to always return a particular type), and room for the implementation to change in the future.
Another example of considering the level of detail to document is when it comes to how input sequences are used in LINQ to Objects. What does it mean to say "this method uses deferred execution" _exactly_? If I call GetEnumerator() eagerly but defer the call to MoveNext(), is that still "deferred execution"? Should the documentation state when a sequence is buffered and when its streamed? Should it guarantee the order of the result sequence when the natural implementation makes that order easy to describe (e.g. for Distinct)? In this series Ive tried to be as clear as possible about what actually happens but thats not to say that in some cases, the documentation wasnt left _deliberately_ ambiguous.
### Conclusion
There are many other design considerations that I havent gone into here particularly optimization, which Ive already covered twice, probably saying everything I wanted to say here anyway.
I may add a few more bits to this post over time… but aside from that, I think Im fundamentally _done_. Ill write one more conclusion post, then declare Edulinq closed…
[C#](https://codeblog.jonskeet.uk/category/csharp/), [Edulinq](https://codeblog.jonskeet.uk/category/edulinq/), [LINQ](https://codeblog.jonskeet.uk/category/linq/)
# [Reimplementing LINQ to Objects: Part 43 Out-of-process queries with IQueryable](https://codeblog.jonskeet.uk/2011/02/20/reimplementing-linq-to-objects-part-43-out-of-process-queries-with-iqueryable/)
[February 20, 2011](https://codeblog.jonskeet.uk/2011/02/20/reimplementing-linq-to-objects-part-43-out-of-process-queries-with-iqueryable/) [jonskeet](https://codeblog.jonskeet.uk/author/jonskeet/) [9 Comments](https://codeblog.jonskeet.uk/2011/02/20/reimplementing-linq-to-objects-part-43-out-of-process-queries-with-iqueryable/#comments)
Ive been putting off writing about this for a while now, mostly because its such a huge topic. Im not going to try to give more than a brief introduction to it here dont expect to be able to whip up your own LINQ to SQL implementation afterwards but its worth at least having an idea of what happens when you use something like LINQ to SQL, NHibernate or the Entity Framework.
Just as LINQ to Objects is primarily interested in IEnumerable<T> and the static Enumerable class, so out-of-process LINQ is primarily interested in IQueryable<T> and the static Queryable class… but before we get to them, we need to talk about expression trees.
### Expression Trees
To put it in a nutshell, expression trees encapsulate logic in _data_ instead of _code_. While you _can_ introspect .NET code via [MethodBase.GetMethodBody](http://msdn.microsoft.com/en-us/library/system.reflection.methodbase.getmethodbody.aspx) and then [MethodBody.GetILAsByteArray](http://msdn.microsoft.com/en-us/library/system.reflection.methodbody.getilasbytearray.aspx), thats not really a practical approach. The types in the [System.Linq.Expressions](http://msdn.microsoft.com/en-us/library/system.linq.expressions.aspx) define expressions in an easier-to-process manner. When expression trees were introduced in .NET 3.5, they were strictly for _expressions_, but the Dynamic Language Runtime uses expression trees to represent operations, and the range of logic represented had to expand accordingly, to include things like blocks.
While you certainly _can_ build expression trees yourself (usually via the factory methods on the nongeneric [Expression](http://msdn.microsoft.com/en-us/library/system.linq.expressions.expression.aspx) class), and its fun to do so at times, the most common way of creating them is to use the C# compilers support for them via lambda expressions. So far weve always seen a lambda expression being converted to a delegate, but it can also convert lambdas to instances of [Expression<TDelegate>](http://msdn.microsoft.com/en-us/library/bb335710.aspx), where TDelegate is a delegate type which is compatible with the lambda expression. A concrete example will help here. The statement:
Expression<Func<int, int>> addOne = x => x + 1;
will be compiled into code which is _effectively_ something like this:
var parameter = Expression.Parameter(typeof(int), "x");
var one = Expression.Constant(1, typeof(int));
var addition = Expression.Add(parameter, one);
var addOne = Expression.Lambda<Func<int, int>>(addition,  new ParameterExpression[] { parameter });
The compiler has some tricks up its sleeves which allow it to refer to methods, events and the like in a simpler way than we can from code, but largely you can regard the transformation as just a way of making life a _lot_ simpler than if you had to build the expression trees yourself every time.
### IQueryable, IQueryable<T> and IQueryProvider
Now that weve got the idea of being able to inspect logic relatively easily at execution time, lets see how it applies to LINQ.
There are three interfaces to introduce, and its probably easiest to start with how they appear in a class diagram:
![IEnumerabIe< Generic Interface -O IEnumerabIe IEnumerabIe Interface IQueryabIe< T > Generic Interface -O IQueryab le -b IEnumerabIe IQ ueryable Interface -O IEnumerabIe Element Type provider Type Abstract Class -D Memberlnfo Exp ression Abstract Class IQ ueryProvider Interface Methods CreoteQuery 0 veriood) Execute overload) ](Exported%20image%2020240808113923-0.png)
Most of the time, queries are represented using the generic [IQueryable<T>](http://msdn.microsoft.com/en-us/library/bb351562.aspx) interface, but this doesnt actually add much over the nongeneric [IQueryable](http://msdn.microsoft.com/en-us/library/system.linq.iqueryable.aspx) interface it extends, other than _also_ extending IEnumerable<T> so you can iterate over the contents of an IQueryable<T> just as with any other sequence.
IQueryable contains the interesting bits, in the form of three properties: [ElementType](http://msdn.microsoft.com/en-us/library/system.linq.iqueryable.elementtype.aspx) which indicates the type of the elements within the query (in other words, a dynamic form of the T from IQueryable<T>), [Expression](http://msdn.microsoft.com/en-us/library/system.linq.iqueryable.expression.aspx) returns the expression tree for the query so far, and [Provider](http://msdn.microsoft.com/en-us/library/system.linq.iqueryable.provider.aspx) returns the query provider which is responsible for creating _new_ queries and executing the existing one. We wont need to use the ElementType property ourselves, but well need both the Provider and Expression properties.
### The static Queryable class
Were not going to implement any of the interfaces ourselves, but Ive got a small sample program to demonstrate how they all work, imagining we were implementing most of [Queryable](http://msdn.microsoft.com/en-us/library/system.linq.queryable.aspx) ourselves. This static class contains extension methods for IQueryable<T> just as Enumerable does for IEnumerable<T>. _Most_ of the query operators from LINQ to Objects appear in Queryable as well, but there are a few notable omissions, such as the To{Lookup, Array, List, Dictionary} methods. If you call one of those on an IQueryable<T>, the Enumerable implementations will be used instead. (IQueryable<T> extends IEnumerable<T>, so the extension methods in Enumerable are applicable to IQueryable<T> sequences as well.)
The big difference between the Queryable and Enumerable methods in terms of their _declarations_ is in the parameters:
- The "source" parameter in Queryable is always of type IQueryable<TSource> instead of IEnumerable<TSource>. (Other sequence parameters such as the sequence to concatenate for Queryable.Concat are expressed as IEnumerable<T>, interestingly enough. This allows you to express a SQL query using "local" data as well; the query methods work out whether the sequence is actually an IQueryable<T> and act accordingly.)
- Any parameters which were delegates in Enumerable are expression trees in Queryable; so while the selector parameter in Enumerable.Select is of type Func<TSource, TResult>, the equivalent in Queryable.Select is of type Expression<Func<TSource, TResult>>
The big difference between the methods in terms of what they _do_ is that whereas the Enumerable methods actually do the work (eventually possibly after deferred execution of course), the Queryable methods themselves really _dont_ do any work: they just ask the query provider to build up a query indicating that theyve been called.
Lets have a look at Where for example. If we wanted to implement Queryable.Where, we would have to:
- Perform argument checking
- Get the "current" querys Expression
- Build a new expression representing a call to Queryable.Where using the current expression as the source, and the predicate expression as the predicate
- Ask the current querys provider to build a new IQueryable<T> based on that call expression, and return it.
It all sounds a bit recursive, I realize the Where call needs to record that a Where call has happened… but thats all. You may very well wonder where all the work is happening. Well come to that.
Now building a call expression is slightly tedious because you need to have the right MethodInfo and as Where is overloaded, that means distinguishing between the two Where methods, which is easier said than done. Ive actually used a LINQ query to find the right overload the one where the predicate parameter Expression<Func<T, bool>> rather than Expression<Func<T, int, bool>>. In the .NET implementation, methods can use [MethodBase.GetCurrentMethod()](http://msdn.microsoft.com/en-us/library/system.reflection.methodbase.getcurrentmethod.aspx) instead… although equally they could have created a bunch of static variables computed at class initialization time. We cant use GetCurrentMethod() for experimentation purposes, because the query provider is likely to expect the exact correct method from System.Linq.Queryable in the System.Core assembly.
Heres our sample implementation, broken up quite a lot to make it easier to understand:
public static IQueryable<TSource> Where<TSource>(
    this IQueryable<TSource> source,
    Expression<Func<TSource, bool>> predicate)
{
    if (source == null)
    {
        throw new ArgumentNullException("source");
    }
    if (predicate == null)
    {
        throw new ArgumentNullException("predicate");
    }
        
    Expression sourceExpression = source.Expression;
    Expression quotedPredicate = Expression.Quote(predicate);
        
    // This gets the "open" method, without specific type arguments. The second parameter
    // of the method we want is of type Expression<Func<TSource, bool>>, so the sole generic
    // type argument to Expression<T> itself has two generic type arguments.
    // Lets face it, reflection on generic methods is a mess.
    MethodInfo method = typeof(Queryable).GetMethods()
                                         .Where(m => m.Name == "Where")
                                         .Where(m => m.GetParameters()[1]
                                                      .ParameterType
                                                      .GetGenericArguments()[0]
                                                      .GetGenericArguments().Length == 2)
                                         .First();
        
    // This gets the method with the same type arguments as ours
    MethodInfo closedMethod = method.MakeGenericMethod(new Type[] { typeof(TSource) });
        
    // Now we can create a *representation* of this exact method call
    Expression methodCall = Expression.Call(closedMethod, sourceExpression, quotedPredicate);
        
    // … and ask our query provider to create a query for it
    return source.Provider.CreateQuery<TSource>(methodCall);
}
Theres only one part of this code that I dont really understand the need for, and thats the call to Expression.Quote on the predicate expression tree. Im sure theres a good reason for it, but _this particular example_ would work without it, as far as I can see. The real implementation uses it though, so dare say its required in some way.
EDIT: Daniels comment has made this somewhat clearer to me. Each of the arguments to Expression.Call after the MethodInfo itself is meant to be an expression which represents the argument to the method call. In our example we need an expression which represents an argument of type Expression<Func<TSource, bool>>. We already have the value, but we need to provide the layer of wrapping… just as we did with Expression.Constant in the very first expression tree I showed at the top. To wrap the expression value weve got, we use Expression.Quote. Its still not clear to me _exactly_ why we can use Expression.Quote but not Expression.Constant, but at least its clearer why we need _something_
EDIT: Im gradually getting there. [This Stack Overflow answer from Eric Lippert](http://stackoverflow.com/questions/3716492/what-does-expression-quote-do-that-expression-constant-cant-already-do) has much to say on the topic. Im still trying to get my head round it, but Im sure when Ive read Erics answer several times, Ill get there.
We can even test that this works, by using the [Queryable.AsQueryable](http://msdn.microsoft.com/en-us/library/bb353734.aspx) method from the real .NET implementation. This creates an IQuerable<T> from any IEnumerable<T> using a built-in query provider. Heres the test program, where FakeQueryable is a static class containing the extension method above:
using System;
using System.Collections.Generic;
using System.Linq;
class Test
{
    static void Main()
    {        
        List<int> list = new List<int> { 3, 5, 1 };
        IQueryable<int> source = list.AsQueryable();
        IQueryable<int> query = FakeQueryable.Where(source, x => x > 2);
        
        foreach (int value in query)
        {
            Console.WriteLine(value);
        }
    }
}
This works, printing just 3 and 5, filtering out the 1. Yay! (Im explicitly calling FakeQueryable.Where rather than letting extension method resolution find it, just to make things clearer.)
Um, but whats doing the actual work? Weve implemented the Where clause without providing any filtering ourselves. Its really the query provider which has built an appropriate IQueryable<T> implementation. When we call GetEnumerator() implicitly in the foreach loop, the query can examine everything thats built up in the expression tree (which could contain multiple operators its nesting queries within queries, essentially) and work out what to do. In the case of our IQueryable<T> built from a list, it just does the filtering in-process… but if we were using LINQ to SQL, _thats_ when the SQL would be generated. The provider recognizes the specific methods from Queryable, and applies filters, projections etc. Thats why it was important that our demo Where method pretended that the real Queryable.Where had been called otherwise the query provider wouldnt know what the call expression
Just to hammer the point home even further… Queryable itself neither knows nor cares what kind of data source youre using. Its job is _not_ to perform any query operations itself; its job is to _record_ the requested query operations in a source-agnostic manner, and let the source provider handle them when it needs to.
### Immediate execution with IQueryProvider.Execute
All the operators using deferred execution in Queryable are implemented in much the same way as our demo Where method. However, that doesnt cover the situation where we need to execute the query _now_, because it has to return a value directly instead of another query.
This time Im going to use ElementAt as the sample, simply because its only got one overload, which makes it very easy to grab the relevant MethodInfo. The general procedure is exactly the same as building a new query, except that this time we call the providers Execute method instead of CreateQuery.
public static TSource ElementAt<TSource>(this IQueryable<TSource> source, int index)
{
    if (source == null)
    {
        throw new ArgumentNullException("source");
    }
        
    Expression sourceExpression = source.Expression;
    Expression indexExpression = Expression.Constant(index);
        
    MethodInfo method = typeof(Queryable).GetMethod("ElementAt");        
    MethodInfo closedMethod = method.MakeGenericMethod(new Type[] { typeof(TSource) });
        
    // Now we can create a *representation* of this exact method call
    Expression methodCall = Expression.Call(closedMethod, sourceExpression, indexExpression);
        
    // … and ask our query provider to execute it
    return source.Provider.Execute<TSource>(methodCall);
}
The type argument we provide to Execute is the desired _return_ type so for Count, wed call Execute<int> for example. Again, its up to the query provider to work out what the call actually means.
Its worth mentioning that both CreateQuery and Execute have generic and non-generic overloads. I havent personally encountered a use for the non-generic ones, but I gather theyre useful for various situations in generated code, particularly if you really dont know the element type or at least only know it dynamically, and dont want to have to use reflection to generate an appropriate generic method call.
### Transparent support in source code
One of the aspects of LINQ which raises it to the "genius" status (and "slightly scary" at the same time) is that most of the time, most developers dont need to make any changes to their source code in order to use Enumerable or Queryable. Take this query expression and its translation:
var query = from person in family
            where person.LastName == "Skeet"
            select person.FirstName;
// Translation
var query = family.Where(person => person.LastName == "Skeet")
                  .Select(person => person.FirstName);
Which set of query methods will that use? It entirely depends on the compile-time type of the "family" variable. If thats a type which implements IQueryable<T>, it will use the extension methods in Queryable, the lambda expression will be converted into expression trees, and the type of "query" will be IQueryable<string>. Otherwise (and assuming the type implements IEnumerable<T> isnt some other interesting type such as [ParallelEnumerable](http://msdn.microsoft.com/en-us/library/system.linq.parallelenumerable.aspx)) it will use the extension methods in Enumerable, the lambda expressions will be converted into delgeates, and the type of "query" will be IEnumerable<string>.
The query expression translation part of the specification has no need to care about this, because its simply translating into a form which uses lambda expressions the rest of overload resolution and lambda expression conversion deals with the details.
Genius… although it does mean you need to be careful that _really_ you know where your query evaluation is going to take place you dont want to accidentally end up performing your whole query in-process having shipped the entire contents of a database across a network connection…
### Conclusion
This was really a whistlestop tour of the "other" side of LINQ and without going into any of the details of the real providers such as LINQ to SQL. However, I hope its given you enough of a flavour for whats going on to appreciate the general design. Highlights:
- _Expression trees_ are used to capture logic in a data structure which can be examined relatively easily at execution time
- Lambda expressions can be converted into expression trees as well as delegates
- IQueryable<T> and IQueryable form a sort of parallel interface hierarchy to IEnumerable<T> and IEnumerable although the queryable forms extend the enumerable forms
- IQueryProvider enables one query to be built based on another, or executed immediately where appropriate
- Queryable provides equivalent extension methods to most of the Enumerable LINQ operators, except that it uses IQueryable<T> sources and expression trees instead of delegates
- Queryable doesnt handle the queries itself at all; it simply records whats been called and delegates the real processing to the query provider
I _think_ Ive now covered most of the topics I wanted to mention after finishing the actual Edulinq implementation. Next up Ill talk about some of the thorny design issues (most of which Ive already mentioned, but which bear repeating) and then Ill write a brief "series conclusion" post with a list of links to all the other parts.
[C#](https://codeblog.jonskeet.uk/category/csharp/), [Edulinq](https://codeblog.jonskeet.uk/category/edulinq/), [LINQ](https://codeblog.jonskeet.uk/category/linq/)
# [Reimplementing LINQ to Objects: Part 42 More optimization](https://codeblog.jonskeet.uk/2011/01/30/reimplementing-linq-to-objects-part-42-more-optimization/)
[January 30, 2011](https://codeblog.jonskeet.uk/2011/01/30/reimplementing-linq-to-objects-part-42-more-optimization/) [jonskeet](https://codeblog.jonskeet.uk/author/jonskeet/) [4 Comments](https://codeblog.jonskeet.uk/2011/01/30/reimplementing-linq-to-objects-part-42-more-optimization/#comments)
A few parts ago, I jotted down a few thoughts on optimization. Three more topics on that general theme have occurred to me, one of them prompted by the comments.
### User-directed optimizations
I mentioned last time that for micro-optimization purposes, we could derive a tiny benefit if there were operators which allowed us to turn off potential optimizations effectively declare in the LINQ query that we believed the input sequence would _never_ be an IList<T> or an ICollection<T>, so it wasnt worth checking it. I still believe that level of optimization would be futile.
However, going the other way is entirely possible. Imagine if we could say, "There are probably a lot of items in this collection, and the operations I want to perform on them are independent and thread-safe. Feel free to parallelize them."
Thats exactly what Parallel LINQ gives you, of course. A simple call to AsParallel() somewhere in the query often at the start, but it doesnt have to be enables parallelism. You need to be careful how you use this, of course, which is why its opt-in… and it gives you a fair amount of control in terms of degrees of potential parallelism, whether the results are required in the original order and so on.
In some ways my "TopBy" proposal is similar in a very small way, in that it gives information relatively early in the query, allowing the subsequent parts (ThenBy clauses) to take account of the extra information provided by the user. On the other hand, the effect is extremely localized basically just for the sequence of clauses to do with ordering.
Related to the idea of parallelism is the idea of side-effects, and how they affect LINQ to Objects itself.
### Side-effects and optimization
The optimizations in LINQ to Objects appear to make some assumptions about side-effects:
- Iterating over a collection wont cause any side-effects
- Predicates _may_ cause side-effects
Without the first point, all kinds of optimizations would effectively be inappropriate. As the simplest example, Count() wont use an iterator it will just take the count of the collection. What if this was an odd collection which mutated something during iteration, though? Or what if accessing the Count property itself had side-effects? At that point wed be violating our principle of not changing observable behaviour by optimizing. Again, the optimizations are basically assuming "sensible" behaviour from collections.
Theres a rather more subtle possible cause of side-effects which Ive never seen discussed. In some situations most obviously Skip an operator can be implemented to move over an iterator for a time without taking each "current" value. This is due to the separation of MoveNext() from Current. What if we were dealing with an iterator which had side-effects _only when Current was fetched_? It would be easy to write such a sequence but again, I suspect theres an implicit assumption that such sequences simply dont exist, or that its reasonable for the behaviour of LINQ operators with respect to them to be left unspecified.
Predicates, on the other hand, might not be so sensible. Suppose we were computing "sequence.Last(x => 10 / x > 1)" on the sequence { 5, 0, 2 }. Iterating over the sequence forwards, we end up with a DivideByZeroException whereas if we detected that the sequence was a list, and worked our way backwards from the end, wed see that 10 / 2 > 1, and return that last element (2) immediately. Of course, exceptions arent the only kind of side-effect that a predicate can have: it _could_ mutate other state. However, its generally easier to spot that and cry foul of it being a proper functional predicate than notice the possibility of an exception.
I believe this is the reason the predicated Last overload _isnt_ optimized. It would be nice if these assumptions were documented, however.
### Assumptions about performance
Theres a final set of assumptions which the common ICollection<T>/IList<T> optimizations have all been making: that using the more "direct" members of the interfaces (specifically Count and the indexer) are more efficient than simply iterating. The interfaces make no such declarations: theres no requirement that Count _has_ to be O(1), for example. Indeed, its not even the case in the BCL. The first time you ask a "view between" on a sorted set for its count after the underlying set has changed, it has to count the elements again.
Ive had this problem before, [removing items from a HashSet in Java](https://codeblog.jonskeet.uk/2010/07/29/there-s-a-hole-in-my-abstraction-dear-liza-dear-liza). The problem is that theres no way of communicating this information in a standardized way. We _could_ use attributes for everything, but it gets very complicated, and I strongly suspect it would be a complete pain to use. Basically, performance is one area where abstractions just dont hold up or rather, the abstractions arent _designed_ to include performance characteristics.
Even if we knew the complexity of (say) Count that still wouldnt help us necessarily. Suppose its an O(n) operation that sounds bad, until you discover that for this particular horrible collection, _each iteration step_ is also O(n) for some reason. Or maybe theres a collection with an O(1) count but a _horrible_ constant value, whereas iterating is really quick per item… so for small values of O(n), iteration would be faster. Then youve got to bear in mind how much processor time is needed trying to work out the fastest approach… its all bonkers.
So instead we make these assumptions, and for the most part theyre correct. Just be aware of their presence.
### Conclusion
I have reached the conclusion that Im tired, and need sleep. I might write about Queryable, IQueryable and query expressions next time.
[C#](https://codeblog.jonskeet.uk/category/csharp/), [Edulinq](https://codeblog.jonskeet.uk/category/edulinq/), [LINQ](https://codeblog.jonskeet.uk/category/linq/)
# [Reimplementing LINQ to Objects: Part 41 How query expressions work](https://codeblog.jonskeet.uk/2011/01/28/reimplementing-linq-to-objects-part-41-how-query-expressions-work/)
[January 28, 2011](https://codeblog.jonskeet.uk/2011/01/28/reimplementing-linq-to-objects-part-41-how-query-expressions-work/) [jonskeet](https://codeblog.jonskeet.uk/author/jonskeet/) [6 Comments](https://codeblog.jonskeet.uk/2011/01/28/reimplementing-linq-to-objects-part-41-how-query-expressions-work/#comments)
Okay, first a quick plug. This _wont_ be in as much detail as chapter 11 of [C# in Depth](http://manning.com/skeet2/). If you want more, buy a copy. (Until Feb 1st, theres 43% off it if you buy it from Manning with coupon code j2543.) Admittedly that chapter has to _also_ explain all the basic operators rather than just query expression translations, but thats a fair chunk of it.
If youre already familiar with query expressions, dont expect to discover anything particularly insightful here. However, you might be interested in the cheat sheet at the end, in case you forget some of the syntax occasionally. (I know I do.)
### What is this "query expression" of which you speak?
Query expressions are a little nugget of goodness hidden away in section 7.16 of the [C# specification](http://csharpindepth.com/Articles/Chapter1/Specifications.aspx). Unlike some language features like generics and dynamic typing, query expressions keep themselves to themselves, and dont impinge on the rest of the spec. A query expression is a bit of C# which looks a bit like a mangled version of SQL. For example:
from person in people
where person.FirstName.StartsWith("J")
orderby person.Age
select person.LastName
It looks somewhat unlike the rest of C#, which is both a blessing and a curse. On the one hand, queries stand out so its easy to see theyre queries. On the other hand… they stand out rather than fitting in with the rest of your code. To be honest I havent found this to be an issue, but it can take a little getting used to.
Every query expression can be represented in C# code, but the reverse isnt true. Query expressions only take in a subset of the standard query operators and only a limited set of the overloads, at that. Its not unusual to see a query expression followed by a "normal" query operator call, for example:
var list = (from person in people
            where person.FirstName.StartsWith("J")
            orderby person.Age
            select person.LastName)
           .ToList();
So, thats very roughly what they look like. Thats the sort of thing Im dealing with in this post. Lets start dissecting them.
### Compiler translations
First its worth introducing the general principle of query expressions: they effectively get translated step by step into C# which eventually _doesnt_ contain any query expressions. To stick with our first example, that ends up be translated into this code:
people.Where(person => person.FirstName.StartsWith("J"))
      .OrderBy(person => person.Age)
      .Select(person => person.LastName)
Its important to understand that the compiler hasnt done anything apart from systematic translation to get to this point. In particular, so far we havent depended on what "people" is, nor "Where", "OrderBy" or "Select".
Can you tell what this code does yet? You can probably hazard a pretty good guess, but you cant tell. Is it going to call Edulinq.Enumerable.Select, or System.Linq.Enumerable.Select, or something entirely different? It depends on the context. Heck, "people" could be the name of a type which has a static Where method. Or maybe it could be a reference to a class which has an _instance_ method called Where… the options are open.
Of course, they dont stay open for long: the compiler takes that expression and compiles it applying all the normal rules. It converts the lambda expression into either a delegate or an expression tree, tries to resolve Where, OrderBy and Select as normal, and life continues. (Dont worry if youre not sure about expression trees yet Ill come to them in another post.)
The important point is that the query expression translations dont know about System.Linq. The spec barely mentioned IEnumerable<T>, and certainly doesnt rely on it. The whole thing is _pattern based_. If you have an API which provides some or all of the operators used by the pattern, in an appropriate way, you can use query expressions with it. Thats the secret sauce that allows you to use the same syntax for LINQ to Objects, LINQ to SQL, LINQ to Entities, Reactive Extensions, Parallel Extensions and more.
### Range variables and the initial "from" clause
The first part of the query to look at is the first "from" clause, at the start of the query. Its worth mentioning upfront that this is handled somewhat differently to any _later_ "from" clauses Ill explain how theyre translated later.
So we have an expression of the form:
from _[type]_ _identifier_ in _expression_
The "expression" part is just any expression. In most cases there _isnt_ a type specified, in which case the translated version is simply the expression, but with the compiler remembering the identifier as a _range variable_. Ill do my best to explain what range variables are in a minute :)
If there _is_ a type specified, that represents a call to Cast<_type_>(). So examples of the two translations so far are:
// Query (incomplete)
from x in people
// Translation (+ range variable of "x")
people
// Query (incomplete)
from Person x in people
// Translation (+ range variable of "x")
(people).Cast<Person>()
These arent complete query expressions queries have very precise rules about how they can start and end. They _always_ start with a "from" clause like this, and always end either with a "group by" clause or a "select" clause.
So whats the point of the range variable? Well, thats what gets used as the name of the lambda expression parameter used in all the later clauses. Lets add a select clause to create a complete expression and demonstrate how the variable could be used.
### A "select" clause
A select clause is usually translated into a call to Select, using the "body" of the clause as the body of the lambda expression… and the range variable as the parameter. So to expand our previous query, we might have this translation:
// Query
from x in people
select x.Name
// Translation
people.Select(x => x.Name)
Thats all that range variables are used for: to provide placeholders within lambda expressions, effectively. Theyre quite unlike normal variables in most senses. It only makes sense to talk about the "value" of a range variable within a particular clause at a particular point in time when the clause is executing, for one particular value. Their nearest conceptual neighbour is probably the iteration variable declared in a foreach statement, but even thats not really the same particularly given the way [iteration variables are captured](http://blogs.msdn.com/b/ericlippert/archive/2009/11/12/closing-over-the-loop-variable-considered-harmful.aspx), often to the surprise of developers.
The body part has to be a single _expression_ you cant use "statement lambdas" in query expressions. For example, theres no query expression which would translate to this:
// Cant express this in a query expression
people.Select(x => { 
                     Console.WriteLine("Got " + x);
                     return x.Name;
                   })
Thats a perfectly valid C# expression, its just theres now way of expressing it directly as a query expression.
I mentioned that a select clause _usually_ translates into a Select call. There are two cases where it doesnt:
- If its the sole clause after a secondary "from" clause, or a "group by", "join" or "join … into" clause, the body is used in the translation of that clause
- If its an "identity" projection coming after another clause, its removed entirely.
Ill deal with the first point when we reach the relevant clauses. The second point leads to these translations:
// Query
from x in people
where x.IsAdult
select x
// Translation: Select is removed
people.Where(x => x.IsAdult)
// Query
from x in people
select x
// Translation: Select is *not* removed
people.Select(x => x)
The point of including the "pointless" select in the second translation is to hide the original source sequence; its assumed that theres no need to do this in the first translation as the "Where" call will already have protected the source sufficiently.
### The "where" clause
This ones really simple especially as weve already seen it! A where clause always just translates into a Where call. Sample translation, this time with no funny business removing degenerate query expressions:
// Query
from x in people
where x.IsAdult
select x.Name
// Translation
people.Where(x => x.IsAdult)
      .Select(x => x.Name)
Note how the range variable is propagated through the query.
### The "orderby" clause
Heres a secret: I can never remember offhand whether its "orderby" or "order by" its confusing because it really is "group by", but "orderby" is actually just a single word. Of course, Visual Studio gives a pretty unsubtle hint in terms of colouring.
In the simplest form, an orderby clause might look like this:
// Query
from x in people
orderby x.Age
select x.Name
// Translation
people.OrderBy(x => x.Age)
      .Select(x => x.Name)
There are two things which can add complexity though:
- You can order by multiple expressions, separating them by commas
- Each expression can be ordered ascending implicitly, ascending _explicitly_ or descending explicitly.
The first sort expression is always translated into OrderBy or OrderByDescending; subsequent ones always become ThenBy or ThenByDescending. It makes no difference whether you explicitly specify "ascending" or not Ive very rarely seen it in real queries. Heres an example putting it all together:
// Query
from x in people
orderby x.Age, x.FirstName descending, x.LastName ascending
select x.LastName
// Translation
people.OrderBy(x => x.Age)
      .ThenByDescending(x => x.FirstName)
      .ThenBy(x => x.LastName)
      .Select(x => x.LastName)
Top tip: dont use multiple "orderby" clauses consecutively. This query is almost certainly _not_ what you want:
// Dont do this!
from x in people 
orderby x.Age
orderby x.FirstName
select x.LastName
That will end up sorting by FirstName and _then_ Age, and doing so rather slowly as it has to sort twice.
### The "group by" clause
Grouping is another alternative to "select" as the final clause in a query. There are two expressions involved: the element selector (what you want to get in each group) and the key selector (how you want the groups to be organized). Unsurprisingly, this uses the GroupBy operator. So you might have a query to group people in families by their last name, with each group containing the first names of the family members:
// Query expression
from x in people 
group x.FirstName by x.LastName
// Translation
people.GroupBy(x => x.LastName, x => x.FirstName)
If the element selector is trivial, it isnt specified as part of the translation:
// Query expression
from x in people 
group x by x.LastName
// Translation
people.GroupBy(x => x.LastName)
### Query continuations
Both "select" and "group by" can be followed by "into _identifier_". This is known as a _query continuation_, and its really simple. Its translation in the specification isnt in terms of a method call, but instead it transforms one query expression into another, effectively nesting one query as the source of another. I find that translation tricky to think about, personally… I prefer to think of it as using a temporary variable, like this:
// Original query
var query = from x in people
            select x.Name into y
            orderby y.Length
            select y[0];
// Query continuation translation
var tmp = from x in people
          select x.Name;
var query = from y in tmp
            orderby y.Length
            select y[0];
// Final translation into methods
var query = people.Select(x => x.Name)
                  .OrderBy(y => y.Length)
                  .Select(y => y[0]);
Obviously that final translation _could_ have been expressed in terms of two statements as well… theyd be equivalent. This is why its important that LINQ uses deferred execution you can split up a query as much as you like, and it wont alter the execution flow. The query wouldnt actually _execute_ when the value is assigned to "tmp" its just _preparing_ the query for execution.
### Transparent identifiers and the "let" clause
The rest of the query expression clauses all introduce an extra range variable in some form or other. This is the part of query expression translation which is hardest to understand, because it affects how any usage of the range variable in the query expression is translated.
Well start with probably the simplest of the remaining clauses: the "let" clause. This simply introduces a new range variable based upon a projection. Its a bit like a "select", but after a "let" clause both the original range variable _and_ the new one are in scope for the rest of the query. Theyre typically used to avoid redundant computations, or simply to make the code simpler to read. For example, suppose computing an employees tax is a complicated operation, and we want to display a list of employees and the tax they pay, with the higher tax-payer first:
from x in employees
let tax = x.ComputeTax()
orderby tax descending
select x.LastName + ": " + tax
Thats pretty readable, and weve managed to avoid computing the tax twice (once for sorting and once for display).
The problem is, both "x" and "tax" are in scope at the same time… so what are we going to pass to the Select method at the end? We need one entity to pass through our query, which knows the value of both "x" and "tax" at any point (after the "let" clause, obviously). This is precisely the point of a transparent identifier. You can think of the above query as being translated into this:
// Translation from "let" clause to another query expression
from x in employees
select new { x, tax = x.ComputeTax() } into z
orderby z.tax descending
select z.x.LastName + ": " + z.tax
// Final translated query
employees.Select(x => new { x, tax = x.ComputeTax() })
         .OrderByDescending(z => z.tax)
         .Select(z => z.x.LastName + ": " + z.tax)
Here "z" is the transparent identifier which Ive made somewhat more opaque by giving it a name. In the specification, the query translations are performed in terms of "*" which clearly isnt a valid identifier, but which stands in for the transparent one.
The good news about transparent identifiers is that most of the time you dont need to think of them at all. They simply let you have multiple range variables in scope at the same time. I find myself only bothering to think about them explicitly when Im trying to work out the full translation of a query expression which uses them. Its worth knowing about them to avoid being stumped by the concept of (say) a select clause being able to use multiple range variables, but thats all.
Now that weve got the basic concept, we can move onto the final few clauses.
### Secondary "from" clauses
Weve seen that the introductory "from" clause isnt actually translated into a method call, but any subsequent ones are. The syntax is still the same, but the translation uses SelectMany. In many cases this is used just like a cross-join (Cartesian product) but its more flexible than that, as the "inner" sequence introduced by the secondary "from" clause can depend on the current value from the "outer" sequence. Heres an example of that. with the call to SelectMany in the translation:
// Query expression
from parent in adults
from child in parent.Children
where child.Gender == Gender.Male
select child.Name + " is a son of " + parent.Name
// Translation (using z for the transparent identifier)
adults.SelectMany(parent => parent.Children,
                  (parent, child) => new { parent, child })
      .Where(z => z.child.Gender == Gender.Male)
      .Select(z => z.child.Name + " is a son of " + z.parent.Name;
Again we can see the effect of the transparent identifier an anonymous type is introduced to propagate the { parent, child } tuple through the rest of the query.
Theres a special case, however if "the rest of the query" is _just_ a "select" clause, we dont need the anonymous type. We can just apply the projection directly in the SelectMany call. Heres a similar example, but this time without the "where" clause:
// Query expression
from parent in adults
from child in parent.Children
select child.Name + " is a child of " + parent.Name
// Translation (using z for the transparent identifier)
adults.SelectMany(parent => parent.Children,
                  (parent, child) => child.Name + " is a child of " + parent.Name)
This same trick is used in GroupJoin and Join, but I wont go into the details there. Its simpler to just provide examples which use the shortcut, instead of including unnecessary extra clauses just to force the transparent identifier to appear in the translation.
Note that just like the introductory "from" clause, you can specify a type for the range variable, which forces a call to "Cast<>".
### Simple "join" clauses (no "into")
A "join" clause without an "into" part corresponds to a call to the Join method, which represents an inner equijoin. In some ways this is like an extra "from" clause with a "where" clause to provide the relevant filtering, but theres a significant difference: while the "from" clause (and SelectMany) allow you to project each element in the outer sequence to an inner sequence, in Join you merely provide the inner sequence directly, once. You also have to specify the two key selectors one for the outer sequence, and one for the inner sequence. The general syntax is:
join _identifier_ in _inner-sequence_ on _outer-key-selector_ equals _inner-key-selector_
The identifier names the extra range variable introduced. Heres an example including the translation:
// Query expression
from customer in customers
join order in orders on customer.Id equals order.CustomerId
select customer.Name + ": " + order.Price
// Translation
customers.Join(orders,
               customer => customer.Id,
               order => order.CustomerId,
               (customer, order) => customer.Name + ": " + order.Price)
Note how if you put the key selectors the wrong way round, its highly unlikely that the result will compile the lambda expression for the outer sequence doesnt "know about" the inner sequence element, and vice versa. The C# compiler is even nice enough to guess the probable cause, and suggest the fix.
### Group joins "join … into"
Group joins look exactly the same as inner joins, except they have an extra "into _identifier_" part at the end. Again, this introduces an extra range variable but its the identifier after the "into" which ends up in scope, not the one after "join"; that one is _only_ used in the key selector. This is easier to see when we look at a sample translation:
// Query expression
from customer in customers
join order in orders on customer.Id equals order.CustomerId into customerOrders
select customer.Name + ": " + customerOrders.Count()
// Translation
customers.GroupJoin(orders,
                    customer => customer.Id,
                    order => order.CustomerId,
                    (customer, customerOrders) => customer.Name + ": " + customerOrders.Count())
If we had tried to refer to "order" in the select clause, the result would have been an error: its not in scope any more. Note that this is _not_ a query continuation unlike "select … into" and "group … into". It introduces a new range variable, but all the previous range variables are still in scope.
Thats it! Thats all the translations that the C# compiler supports. VBs query expressions are rather richer but I suspect thats at least _partly_ because its more painful to write the "dot notation" syntax in VB, as the lambda expression syntax isnt as nice as C#s.
### Translation cheat sheet
I thought it would be useful to produce a short table of the kinds of clauses supported in query expressions, with the translation used by the C# compiler. The translation is given assuming a single range variable named "x" is in scope. I havent given the alternative options where transparent identifiers are introduced this table isnt meant to be a replacement for all the information above! (Likewise this doesnt mention the optimizations for degenerate query expressions or "identity projection" groupings.)
| | |
|---|---|
|**Query expression clause**|**Translation**|
|First "from _[type]_ _x_ in _sequence_"|Just "sequence" or "sequence.Cast<type>()", but with the introduction of a range variable|
|Subsequent "from" clauses: <br>"from _[type] y_ in _projection_"|SelectMany(x => projection, (x, y) => new { x, y }) <br>or SelectMany(x => projection.Cast<type>(), (x, y) => new { x, y })|
|where _predicate_|Where(x => predicate)|
|select _projection_|Select(x => projection)|
|let y = projection|Select(x => new { x, y = projection })|
|orderby o1, o2 ascending, o3 descending <br>(Each ordering may have descending or ascending specified explicitly; the default is ascending)|OrderBy(x => o1) <br>.ThenBy(x => o2) <br>.ThenByDescending(x => o3)|
|group _projection_ by _key-selector_|GroupBy(x => key-selector, x => projection)|
|join _y_ in _inner-sequece_ <br>_on_ _outer-key-selector_ equals _inner-key-selector_|Join(x => outer-key-selector, <br>    y => inner-key-selector, <br>    (x, y) => new { x, y })|
|join _y_ in _inner-sequece_ <br>_on_ _outer-key-selector_ equals _inner-key-selector_ <br>_into_ _z_|GroupJoin(x => outer-key-selector, <br>    y => inner-key-selector, <br>    (x, z) => new { x, z })|
|_query1_ into _y_ <br>_query2_|(Translation in terms of a new query expression) <br>from y in (query1) <br>query2|
### Conclusion
Hopefully thats made a certain amount of sense out of a fairly complicated topic. I find its one of those "aha!" things at some point it clicks, and then seems reasonably simple (aside from transparent identifiers, perhaps). Until that time, query expressions can be a bit magical.
As an aside, I have a sneaking suspicion that one of my first blog posts consisted of my initial impressions of LINQ, written in a plane on the way to the MVP conference in Seattle in September 2005. I would check, but Im finishing this post in another plane, this time on the way to San Francisco. I think Id have been somewhat surprised to be told in 2005 that Id still be writing blog posts about LINQ over five years later. Mind you, I can think of any number of things which have happened in the intervening years which would have astonished me to about the same degree.
Next time: some more thoughts on optimization. Oh, and Im likely to update my wishlist of extra operators as well, but within the existing post.
[C#](https://codeblog.jonskeet.uk/category/csharp/), [Edulinq](https://codeblog.jonskeet.uk/category/edulinq/), [LINQ](https://codeblog.jonskeet.uk/category/linq/)
# [Reimplementing LINQ to Objects: Part 40 Optimization](https://codeblog.jonskeet.uk/2011/01/26/reimplementing-linq-to-objects-part-40-optimization/)
[January 26, 2011](https://codeblog.jonskeet.uk/2011/01/26/reimplementing-linq-to-objects-part-40-optimization/) [jonskeet](https://codeblog.jonskeet.uk/author/jonskeet/) [8 Comments](https://codeblog.jonskeet.uk/2011/01/26/reimplementing-linq-to-objects-part-40-optimization/#comments)
Im not an expert in optimization, and most importantly I dont have any real-world benchmarks to support this post, so please take it with a pinch of salt. That said, lets dive into what optimizations are available in LINQ to Objects.
### What do we mean by optimization?
Just as we think of refactoring as changing the internal structure of code without changing its externally visible behaviour, optimization is the art/craft/science/voodoo of changing the _performance_ of code without changing its externally visible behaviour. Sort of.
This requires two definitions: "performance" and "externally visible behaviour". Neither are as simple as they sound. In almost all cases, performance is a trade-off, whether in speed vs memory, throughput vs latency, big-O complexity vs the factors _within_ that complexity bound, and best case vs worst case.
In LINQ to Objects the "best case vs worst case" is the balance which we need to consider most often: in the cases where we can make a saving, how significant is that saving? How much does it cost in _every_ case to make a saving _some_ of the time? How often do we actually take the fast path?
Externally visible behaviour is even harder to pin down, because we definitely dont mean _all_ externally visible behaviour. Almost all the optimizations in Edulinq are visible if you work hard enough indeed, thats how I have unit tests for them. I can test that Count() uses the Count property of an ICollection<T> instead of iterating over it by creating an ICollection<T> implementation which works with Count but throws an exception if you try to iterate over it. I dont think we care about that sort of change to externally visible behaviour.
What we really mean is, "If we use the code in a _sensible_ way, will we get the same results with the optimized code as we would without the optimization?" Much better. Nothing woolly about the term "sensible" at all, is there? More realistically, we could talk about a system where every type adheres to the contracts of every interface it implements that would at least get rid of the examples used for unit testing. Still, even the performance of a system is externally visible… its easy to tell the difference between an implementation of Count() which is optimized and one which isnt, if youve got a list of 10 million items.
### How can we optimize in LINQ to Objects?
Effectively we have one technique for significant optimization in LINQ to Objects: finding out that a sequence implements a more capable interface than IEnumerable<T>, and then using that interface. (Or in the case of my optimization for HashSet<T> and Contains, using a concrete type in the same way.) Count() is the most obvious example of this: if the sequence implements ICollection or ICollection<T>, then we can use the Count property on that interface and were done. No need to iterate at all.
These are generally good optimizations because they allow us to transform an O(n) computation into an O(1) computation. Thats a pretty big win when its applicable, and the cost of checking is _reasonably_ small. So long as we hit the right type once every so often and particularly if in those cases the sequences are long then its a net gain. The performance characteristics are likely to be reasonably fixed here for any one program its not like well _sometimes_ win and _sometimes_ lose for a specific query… its that for some queries we win and some we wont. If we were micro-optimizing, we might want a way of calling a "non-optimized" version which didnt even bother trying the optimization, because we know it will always fail. I would regard such an approach as a colossal waste of effort in the majority of cases.
Some optimizations are slightly less obvious, particularly because they dont offer a change in the big-O complexity, but can still make a significant difference. Take ToArray, for example. If we know the sequence is an IList<T> we can construct an array of exactly the right size and ask the list to copy the elements into it. Chances are that copy can be very efficient indeed, basically copying a whole block of bits from one place to another and we know we wont need any resizing. Compare that with building up a buffer, resizing periodically including copying all the elements weve already discovered. Every part of that process is going to be slower, but theyre both O(n) operations really. This is a good example of where big-O notation doesnt tell the whole story. Again, the optimization is almost certainly a good one to make.
Then there are distinctly dodgy optimizations which _can_ make a difference, but are unlikely to apply. My optimization for ElementAt and ElementAtOrDefault comes into play here. Its fine to check whether an object implements IList<T>, and use the indexer if so. Thats an obvious win. But I have an extra optimization to exit quickly if we can find out that the given index is out of the bounds of the sequence. Unfortunately that optimization is only useful when:
- The sequence implements ICollection<T> or ICollection (but remember it has to implement IEnumerable<T> there arent many collections implementing _only_ the non-generic ICollection, but the generic IEnumerable<T>)
- The sequence _doesnt_ implement IList<T> (which gets rid of almost all implementations of ICollection<T>)
- The given index is _actually_ greater than or equal to the size of the collection
All that comes at the cost of a couple of type checks… not a great cost, and we _do_ potentially save an O(n) check for being given an index out of the bounds of the collection… but how often are we really going to make that win? This is where Id love to have something like [Dapper](http://research.google.com/pubs/pub36356.html), but applied to LINQ to Objects and running in a significant number of real-world projects, just logging in as light a way as possible how often we win, how often we lose, and how big the benefit is.
Finally, we come to the optimizations which dont make sense to me… such as the optimization for First in both Mono and LinqBridge. Both of these projects check whether the sequence is a list, so that they check the count and then use the indexer to fetch item 0 instead of calling GetEnumerator()/MoveNext()/Current. Now yes, theres a chance this avoids creating an extra object (although not always, [as weve seen before](https://codeblog.jonskeet.uk/2011/01/18/gotcha-around-iterator-blocks)) but theyre both O(1) operations which are likely to be darned fast. At this point not only is the payback very small (if it even exists) but the whole operation is likely to be so fast that the tiny check for whether the object implements IList<T> is likely to become more significant. Oh, and then theres the extra code complexity yes, thats only relevant to the implementers, but Id personally rather they spent their time on other things (like getting OrderByDescending to work properly… smirk). In other words, I think this is a _bad_ target for optimization. At some point Ill try to do a quick analysis of just how often the collection has to implement IList<T> in order for it to be worth doing this and whether the improvement is even measurable.
Of course there are other micro-optimizations available. When we dont need to fetch the current item (e.g. when skipping over items) lets just call MoveNext() instead of also assigning the return value of a property to a variable. Ive done that in various places in Edulinq, but _not_ as an optimization strategy, which I suspect wont make a significant difference, but for readability to make it clearer to the reader that were just moving along the iterator, not examining the contents as we go.
The only other piece of optimization I think Ive performed in Edulinq is the "yield the first results before sorting the rest" part of my quicksort implementation. Im reasonably proud of that, at least conceptually. I dont think it really fits into any other bucket its just a matter of thinking about what we really need and when, deferring work just in case we never need to do it.
### What can we _not_ optimize in LINQ to Objects?
Ive found a few optimizations in both Edulinq and other implementations which I believe to be invalid.
Heres an example I happened to look at just this morning, when reviewing the code for [Skip](https://codeblog.jonskeet.uk/2011/01/02/reimplementing-linq-to-objects-part-23-take-skip-takewhile-skipwhile):
var list = source as IList<TSource>;
if (list != null)
{
    count = Math.Max(count, 0);
    // Note that "count" is the count of items to skip
    for (int index = count; index < list.Count; index++)
    {
        yield return list[index];
    }
    yield break;
}
If our sequence is a list, we can just skip straight to the right part of it and yield the items one at a time. That sounds great, but what if the list changes (or is even truncated!) while were iterating over it? An implementation working with the simple iterator would usually throw an exception, as the change would invalidate the iterator. This is definitely a behavioural change. When I first wrote about Skip, I included this as a "possible" optimization and actually turned it on in the Edulinq source code. I now believe it to be a mistake, and have removed it completely.
Another example is Reverse, and how it should behave. The documentation is fairly unclear, but when I ran the tests, the Mono implementation used an optimization whereby if the sequence is a list, it will just return items from the tail end using the indexer. (This has now been fixed the Mono team is quick like that!) Again, that means that changes made to the list while iterating will be reflected in the reversed sequence. I believe the documentation for Reverse _ought_ to be clear that:
- Execution is deferred: the input sequence isnt read when the method is called.
- When the result sequence is first read by the caller, a snapshot is taken, and _thats_ whats used to return the data.
- If the result sequence is read more than once (i.e. GetEnumerator is called more than once) then a new snapshot is created each time so changes to the input sequence between calls to GetEnumerator on the result sequence _will_ be observed.
Now this is still not as precise as it might be in terms of what "reading" a sequence entails in particular, a simple implementation of Reverse (as per Edulinq) will actually take the snapshot on the first call to MoveNext() on the iterator returned by GetEnumerator() but thats probably not too bad. The snapshotting behaviour itself is important though, and should be made explicit in my opinion.
The problem with both of these "optimizations" is arguably that theyre applying list-based optimizations _within an iterator block used for deferred execution_. Optimizing for lists either upfront at the point of the initial method call or within an immediate execution operator (Count, ToList etc) is fine, because we assume the sequence wont change during the course of the methods execution. We cant make that assumption with an iterator block, because the flow of the code is very different: our code is visited repeatedly based on the callers use of MoveNext().
##### Sequence identity
Another aspect of behaviour which isnt well-specified is that of identity. When is it valid for an operator to return the input sequence itself as the result sequence?
In the Microsoft implementation, this can occur in two operators: AsEnumerable (which _always_ returns the input sequence reference, even if its null) and Cast (which returns the original reference only if it actually implements IEnumerable<TResult>).
In Edulinq, I have two other operators which can return the input sequence: OfType (only if the original reference implements IEnumerable<TResult> and TResult is a non-nullable value type) and Skip (if you provide a count which is zero or negative). Are these valid optimizations? Lets think about why we might _not_ want them to be…
If youre returning a sequence from one layer of your code to another, you usually want that sequence to be viewed _only_ as a sequence. In particular, if its backed by a List<T>, you dont want callers casting to List<T> and modifying the list. With any operator implemented by an iterator block, thats fine the object returned from the operator has no accessible reference to its input, and the type itself only implements IEnumerable<T> (and IEnumerator<T>, and IDisposable, etc but not IList<T>). Its not so good if the operator decides its okay to return the original reference.
The C# language specification refers to this in the section about query expression translation: a no-op projection at the end of a query can be omitted _if and only if there are other operators in the query_. So a query expression of "from foo in bar select foo" will translate to "bar.Select(foo => foo)" but if we had a "where" clause in the query, the Select call would be removed. Its worth noting that the call to "Cast" generated when you explicitly specify the type of a range variable is _not_ enough to prevent the "no-op" projection from being generated… its almost as if the C# team "knows" that Cast can leak sequence identity whereas Where cant.
Personally I think that the "hiding" of the input sequence should be guaranteed where it makes sense to do so, and explicitly _not_ guaranteed otherwise. We could also add an operator of something like "HideIdentity" which would simply (and unconditionally) add an extra iterator block into the pipeline. That way library authors wouldnt have to guess, and would have a clear way of expressing their intention. Using Select(x => x) or Skip(0) is _not_ clear, and in the case of Skip it would even be pointless when using Edulinq.
As for whether my optimizations are valid thats up for debate, really. It seems hard to justify why leaking sequence identity would be okay for Cast but _not_ okay for OfType, whereas I think theres a better case for claiming that Skip should always hide sequence identity.
##### The Contains issue…
If you remember, I have a disagreement around what Contains should do when you dont provide an equality comparer, and when the sequence implements ICollection<T>. I believe it should be consistent with the rest of LINQ to Objects, which _always_ uses the default equality comparer for the element type when it needs one but the user hasnt specified one. Everyone else (Microsoft, Mono, LinqBridge) has gone with delegating to the collections implementation of ICollection<T>.Contains. That plays well in terms of consistency of what happens if you call Contains on that object, so that it doesnt matter what the compile-time type is. Thats a debate to go into in another post, but I just want to point out that this is _not_ an example of optimization. In some cases it may be faster (notably for HashSet<T>) but it stands a _very_ good chance of changing the behaviour. There is absolutely nothing to suggest that the equality comparer used by ICollection<T> should be the default one for the type and in some cases it definitely isnt.
Its therefore a matter of what _result_ we want to get, not how to get that result faster. Its correctness, not optimization but both the LinqBridge and Mono tests which fail for Edulinq are called "Contains_CollectionOptimization_ReturnsTrueWithoutEnumerating" and I think that shows a mistaken way of thinking about this.
### Can we go further?
Ive been considering a couple of optimizations which I believe to be perfectly legitimate, but which none of the implementations Ive seen have used. One reason I havent implemented them myself yet is that they will reduce the effectiveness of all my unit tests. You see, Ive generally used Enumerable.Range as a good way of testing a non-list-based sequence… but whats to stop Range and Repeat being implemented _as_ IList<T> implementations?
All the non-mutating members are easy to implement, and we can just throw exceptions from the mutating members (as other read-only collections do).
Would this be more efficient? Well yes, if you ever performed a Count(), ElementAt(), ToArray(), ToList() etc operation on a range or a repeated element… but how often is _that_ going to happen? I suspect its pretty rare probably rare enough not to make it worth my time, particularly when you then consider all the tests that would have to be rewritten to use something other than Range when I wanted a non-list sequence…
### Conclusion
Surprise, surprise doing optimization well is difficult. When its obvious what _can_ be done, its not obvious what _should_ be done… and sometimes its not even what is valid in the first place.
Note that none of this has really talked about data structures and algorithms. I looked at some options when implementing ordering, and Im _still_ thinking about the best approach for implementing TopBy (probably either a heap or a self-balancing tree something which could take advantage of the size being constant would be nice) but _in general_ the optimizations here havent required any cunning knowledge of computer science. Thats quite a good thing, because its many years since Ive studied CS seriously…
I suspect that with this post more than almost any other, Im likely to want to add extra items in the future (or amend mistakes which reveal my incompetence). Watch this space.
Next up, I think it would be worth revisiting query expressions from scratch. Anyone whos read C# in Depth or has followed this blog for long enough is likely to be able to skip it, but I think the series would be incomplete without a quick dive into the compiler translations involved.
[C#](https://codeblog.jonskeet.uk/category/csharp/), [Edulinq](https://codeblog.jonskeet.uk/category/edulinq/), [LINQ](https://codeblog.jonskeet.uk/category/linq/)
# [Reimplementing LINQ to Objects: Part 39 Comparing implementations](https://codeblog.jonskeet.uk/2011/01/25/reimplementing-linq-to-objects-part-39-comparing-implementations/)
[January 25, 2011](https://codeblog.jonskeet.uk/2011/01/25/reimplementing-linq-to-objects-part-39-comparing-implementations/) [jonskeet](https://codeblog.jonskeet.uk/author/jonskeet/) [12 Comments](https://codeblog.jonskeet.uk/2011/01/25/reimplementing-linq-to-objects-part-39-comparing-implementations/#comments)
While implementing Edulinq, I only focused on two implementations: .NET 4.0 and Edulinq. However, I was aware that there were other implementations available, notably [LinqBridge](http://linqbridge.googlecode.com/) and the one which comes with [Mono](http://mono-project.com/). Obviously its interesting to see how other implementations behave, so Ive now made a few changes in order to make the test code run in these different environments.
### The test environments
Im using Mono 2.8 (I cant remember the minor version number offhand) but I tend to think of it as "Mono 3.5" or "Mono 4.0" depending on which runtime Im using and which base libraries Im compiling against, to correspond with the .NET versions. Both runtimes ship as part of Mono 2.8. I will use these version numbers for this post, and ask forgiveness for my lack of precision: whenever you see "Mono 3.5" please just think "Mono 2.8 running against the 2.0 runtime, possibly using some of the class libraries normally associated with .NET 3.5".
LinqBridge is a bit like Edulinq a clean room implementation of LINQ to Objects, but built against .NET 2.0. It contains its own Func delegate declarations and its own version of ExtensionAttribute for extension methods. In my experience this makes it difficult to use with the "real" .NET 3.5, so my build targets .NET 2.0 when running against LinqBridge. This means that tests using HashSet had to be disabled. The version of LinqBridge Im running against is 1.2 the latest binary available on the web site. This has AsEnumerable as a plain static method rather than an extension method; the code has been fixed in source control, but I wanted to run against a prebuilt binary, so Ive just disabled my own AsEnumerable tests for LinqBridge. Likewise the tests for Zip are disabled both for LinqBridge and the "Mono 3.5" tests as Zip was only introduced in .NET 4.
The other issue of not having .NET 4 available in the tests is that the string.Join<T>(string, IEnumerable<T>) overload is unavailable something Id used quite a lot in the test code. Ive created a new static class called "StringEx" and replaced string.Join with StringEx.Join everywhere.
There are batch files under a new "testing" directory which will build and run:
- Microsofts LINQ to Objects and Edulinq under .NET
- LinqBridge, Mono 3.5s LINQ to Objects and Edulinq under Mono 3.5
- Mono 4.0s LINQ to Objects and Edulinq under Mono 4.0
Although I have LinqBridge running under .NET 2.0 in Visual Studio, its a bit of a pain building the tests from a batch file (at least without just calling msbuild). The failures running under Mono 3.5 are the same as those running under .NET 2.0 as far as I can tell, so Im not too worried.
Note that while I have built the Mono tests under both the 3.5 and 4.0 profiles, the results were the same other than due to generic variance, so Ive only included the results of the 4.0 profile below.
### What do the tests cover?
Dont forget that the Edulinq tests were written in the spirit of investigation. They cover aspects of LINQs behaviour which are not guaranteed, both in terms of optimization and simple correctness of behaviour. I have included a test which demonstrates the "issue" with calling Contains on an ICollection<T> which uses a non-default equality comparer, as well as the known issue with OrderByDescending using a comparer which returns int.MinValue. There are optimizations which are present in Edulinq but not in LINQ to Objects, and I have tests for those, too.
The tests which fail against Microsofts implementation (for known reasons) are normally marked with an [Ignore] attribute to prevent them from alarming me unduly during development. NUnit categories would make more sense here, but I dont believe ReSharper supports them, and thats the way I run the tests normally. Likewise the tests which take a very long time (such as counting more than int.MaxValue elements) are normally suppressed.
In order to truly run _all_ my tests, I now have a horrible hack using conditional compilation: if the ALL_TESTS preprocessor symbol is defined, I build my own IgnoreAttribute class in the Edulinq.Tests namespace, which effectively takes precedence over the NUnit one… so NUnit will ignore the [Ignore], so to speak. Frankly all this conditional compilation is pretty horrible, and I wouldnt use it for a "real" project, but this is a slightly unusual situation.
EDIT: It turns out that ReSharper _does_ support categories. Im not sure how far that support goes yet, but at the very least theres "Group by categories" available. I may go through _all_ my tests and apply a category to each one: optimization, execution mode, time-consuming etc. Well see whether I can find the energy for that :)
So, lets have a look at what the test results are…
### Edulinq
Unsurprisingly, Edulinq passes all its own tests, with the minor exception of CastTest.OriginalSourceReturnedDueToGenericCovariance running under Mono 3.5, which doesnt include covariance. Arguably this test should be conditionalised to not even run in that situation, as its not expected to work.
### Microsofts LINQ to Objects
8 failures, all expected:
- Contains delegates to the ICollection<T>.Contains implementation if it exists, rather than using the default comparer for the type. This is a design and documentation issue which Ive discussed in more detail in the [Contains part of this series](https://codeblog.jonskeet.uk/2011/01/12/reimplementing-linq-to-objects-part-32-contains).
- Optimization: ElementAt and ElementAtOrDefault dont validate the specified index eagerly when the input sequence implements ICollection<T> but not IList<T>.
- Optimization: OfType always uses an intermediate iterator even when the input sequence already implements IEnumerable<T> and T is a non-nullable value type.
- Optimization: SequenceEqual doesnt compare the counts of the sequences eagerly even when both sequences implement ICollection<T>
- Correctness: OrderByDescending doesnt work if you use a key comparer which returns int.MinValue
- Consistency: Single and SingleOrDefault (with a predicate) dont throw InvalidOperationException as soon as they encounter a second element matching the predicate; the predicate-less overloads _do_ throw as soon as they see a second element.
All of these have been discussed already, so I wont go into them now.
### LinqBridge
LinqBridge had a total of 33 failures. I havent looked into them in detail, but just going from the test output Ive broken them down into the following broad categories:
- Optimization:
- Cast never returns the original source, presumably always introducing an intermediate iterator.
- All three of Microsofts "missed opportunities" listed above are also missed in LinqBridge
- Use of input sequences:
- Except and Intersect appear to read the first sequence first (possibly completely?) and then the second sequence. Edulinq and LINQ to Objects read the second sequence completely and then stream the first sequence. This behaviour is undocumented.
- Join, GroupBy and GroupJoin appear not to be deferred at all. If Im right, this is a definite bug.
- Aggregation accuracy: both Average and Sum over an IEnumerable<float> appear to use a float accumulator instead of a double. This is probably worth fixing for the sake of both range and accuracy, but isnt specified in the documentation.
- OrderBy (etc) appears to apply the key selector multiple times while sorting. The behaviour here isnt documented, but as I [mentioned before](https://codeblog.jonskeet.uk/2011/01/05/reimplementing-linq-to-objects-part-26b-orderby-descending-thenby-descending), it could produce performance issues unnecessarily.
- Exceptions:
- ToDictionary should throw an exception if you give it duplicate keys; it appears not to at least when a custom comparer is used. (Its possible its just not passing the comparer along.)
- The generic Max and Min methods dont return the null value for the element type when that type is nullable. Instead, they throw an exception which is the normal behaviour if the element type is non-nullable. This behaviour isnt well documented, but is consistent with the behaviour of the non-generic overloads. See the [Min/Max post](https://codeblog.jonskeet.uk/2011/01/09/reimplementing-linq-to-objects-part-29-min-max) for more details.
- General bugs:
- The generic form of Min/Max appears not to ignore null values when the element type is nullable.
- OrderByDescending appears to be broken in the same way as Microsofts implementation
- Range appears to be broken around its boundary testing.
- Join, GroupJoin, GroupBy and ToLookup break when presented with null keys
### Mono 4.0 (and 3.5, effectively)
Mono failed 18 of the tests. There are fewer definite bugs than in LinqBridge, but its definitely not perfect. Heres the breakdown:
- Optimization:
- Mono misses the same three opportunities that LinqBridge and Microsoft miss.
- Contains(item) delegates to ICollection<T> when its implemented, just like in the Microsoft implementation. (I assume the authors would call this an "optimization", hence its location in this section.) I believe that LinqBridge has the same behaviour, but that test didnt run in the LinqBridge configuration as it uses HashSet.
- Average/Sum accumulator types:
- Mono appears to use float when working with float values, leading to more accumulator error than is necessary.
- Average overflow for integer types
- Mono appears to use checked arithmetic when _summing_ a sequence, but not when taking the _average_ of a sequence. So the average of { long.MaxValue, long.MaxValue, 2 } is 0. (This originally confused me into thinking it was using floating point types during the summation, but I now believe its just a checked/unchecked issue.)
- Bugs:
- Count doesnt overflow either with or without a predicate
- The Max handling of double.NaN isnt in line with .NET. I havent investigated the reason for this yet.
- OrderByDescending is broken in the same way as for LinqBridge and the Microsoft implementation.
- Range is broken for both Range(int.MinValue, 0) and Range(int.MaxValue, 1). Test those boundary cases, folks :)
- When reversing a list, Mono _doesnt_ buffer the current contents. In other words, changes made while iterating over the reversed list are visible in the returned sequence. The documentation isnt very clear about the desired behaviour here, admittedly.
- GroupJoin and Join match null keys, unlike Microsofts implementation.
### How does Edulinq fare against other unit tests?
It didnt seem fair to _only_ test other implementations against the Edulinq tests. After all, its only natural that my tests should work against my own code. What happens if we run the Mono and LinqBridge tests against my code?
The LinqBridge tests didnt find anything surprising. There were two failures:
- I dont have the "delegate Contains to ICollection<T>.Contains" behaviour, which the tests check for.
- I dont optimize First in the case of the collection implementing IList<T>. I view this as a pretty dubious optimization to be honest I doubt that creating an iterator to get to the first item is going to be much slower than checking for IList<T>, fetching the count, and then fetching the first item via the indexer… and it means that all _non-list_ implementations also have to check whether the sequence implements IList<T>. I dont intend to change Edulinq for this.
The Mono tests picked up the same two failures as above, and two _genuine_ bugs:
- By implementing Take via TakeWhile, I was iterating too far: in order for the condition to become false, we had to iterate to the first item we _wouldnt_ return.
- ToLookup didnt accept null keys a fault which propagated to GroupJoin, Join and GroupBy too. (EDIT: It turns out that its more subtle than that. Nothing should break, but the MS implementation _ignores_ null keys for Join and GroupJoin. Edulinq now does the same, but Ive raised a [Connect issue](https://connect.microsoft.com/VisualStudio/feedback/details/639943/enumerable-join-groupjoin-ignores-null-keys-inconsistency-with-tolookup-and-groupby) to suggest this should _at least_ be documented.)
Ive fixed these in source control, and will add an addendum to each of the relevant posts ([Take](https://codeblog.jonskeet.uk/2011/01/02/reimplementing-linq-to-objects-part-23-take-skip-takewhile-skipwhile), [ToLookup](https://codeblog.jonskeet.uk/2010/12/31/reimplementing-linq-to-objects-part-18-tolookup)) when I have a moment spare.
Theres one additional failure, trying to find the average of a sequence of two Int64.MaxValue values. That overflows on both Edulinq and LINQ to Objects thats the downside of using an Int64 to sum the values. As mentioned, Mono suffers a degree of inaccuracy instead; its all a matter of trade-offs. (A _really_ smart implementation might use Int64 while possible, and then go up to using Double where necessary, I suppose.)
Unfortunately I dont have the tests for the Microsoft implementation, of course… Id love to know whether theres anything Ive failed with there.
### Conclusion
This was very interesting theres a mixture of failure conditions around, and plenty of "non-failures" where each implementations tests are enforcing their own behaviour.
I do find it amusing that all three of the "mainstream" implementations have the same OrderByDescending bug though. Other than that, the clear bugs between Mono and LinqBridge dont intersect, which is slightly surprising.
Its nice to see that despite not setting out to create a "production-quality" implementation of LINQ to Objects, thats _mostly_ what Ive ended up with. Who knows maybe some aspects of my implementation or tests will end up in Mono in the future :)
Given the various different optimizations mentioned in this post, I think its only fitting that next time Ill discuss where we _can_ optimize, where its _worth_ optimizing, and some more tricks we could still pull out of the bag…
[C#](https://codeblog.jonskeet.uk/category/csharp/), [Edulinq](https://codeblog.jonskeet.uk/category/edulinq/), [LINQ](https://codeblog.jonskeet.uk/category/linq/)
# [Reimplementing LINQ to Objects: Part 38 Whats missing?](https://codeblog.jonskeet.uk/2011/01/22/reimplementing-linq-to-objects-part-38-what-s-missing/)
[January 22, 2011](https://codeblog.jonskeet.uk/2011/01/22/reimplementing-linq-to-objects-part-38-what-s-missing/) [jonskeet](https://codeblog.jonskeet.uk/author/jonskeet/) [25 Comments](https://codeblog.jonskeet.uk/2011/01/22/reimplementing-linq-to-objects-part-38-what-s-missing/#comments)
I mentioned before that the Zip operator was only introduced in .NET 4, so clearly theres a _little_ wiggle room for LINQ to Objects query operators to grow in number. This post mentions some of the ones I think are most sorely lack either because Ive wanted them myself, or because Ive seen folks on Stack Overflow want them for entirely reasonable use cases.
There is an issue with respect to other LINQ providers, of course: as soon as some useful operators are available for LINQ to Objects, there will be people who want to apply them to LINQ to SQL, the Entity Framework and the like. Worse, if theyre _not_ included in Queryable with overloads based on expression trees, the LINQ to Objects implementation will silently get picked leading to what looks like a lovely query performing like treacle while the client slurps over the entire database. If they _are_ included in Queryable, then third party LINQ providers could end up with a nasty versioning problem. In other words, some care is needed and Im glad Im not the one who has to decide how new features are introduced.
Ive deliberately _not_ looked at the extra set of operators introduced in the System.Interactive part of [Reactive Extensions](http://msdn.microsoft.com/en-us/devlabs/ee794896)… nor have I looked back over what weve implemented in [MoreLINQ](http://morelinq.googlecode.com/) (an open source project I started specifically to create new operators). I figured it would be worth thinking about this afresh but look at both of those projects for actual implementations instead of just ideas.
Currently theres no implementation of any of this in Edulinq but I could potentially create an "Edulinq.Extras" assembly which made it all available. Let me know if any of these sounds particularly interesting to see in terms of implementation.
### FooBy
I _love_ OrderBy and ThenBy, with their descending cousins. Theyre so much cleaner than building a custom comparer which just performs a comparison between two properties. So why stop with ordering? Theres a whole bunch of operators which could do with some "FooBy" love. For example, imagine we have a list of files, and we want to find the longest one. We dont want to perform a total ordering by size descending, nor do we want to find the maximum file size itself: we want the file _with_ the maximum size. Id like to be able to write that query as:
FileInfo biggestFile = files.MaxBy(file => file.Length);
Note that we can get a _similar_ result by performing one pass to find the maximum length, and then another pass to find the file with that length. However, thats inefficient and assumes we can read the sequence twice (and get the same results both times). Theres no need for that. We could get the same result using Aggregate with a pretty complicated aggregation, but I think this is a sufficiently common case to deserve its own operator.
Wed want to specify which value would be returned if multiple files had the same length (my suggestion would be the first one we encountered with that length) and we could also specify a key comparer to use. The signatures would look like this:
public static TSource MaxBy<TSource, TKey>(
    this IEnumerable<TSource> source,
    Func<TSource, TKey> keySelector)
public static TSource MaxBy<TSource, TKey>(
    this IEnumerable<TSource> source,
    Func<TSource, TKey> keySelector,
    IComparer<TKey> comparer)
Now its not just Max and Min that gain from this "By" idea. It would be useful to apply the same idea to the set operators. The simplest of these to think about would be DistinctBy, but UnionBy, IntersectBy and ExceptBy would be reasonable too. In the case of ExceptBy and IntersectBy we _could_ potentially take the key collection to indicate the keys of the elements we wanted to exclude/include, but it would probably be more consistent to force the two input sequences to be of the same type (as they would have to be for UnionBy and IntersectBy of course). ContainsBy _might_ be useful, but that would effectively be a Select followed by a normal Contains possibly not useful enough to merit its own operator.
### TopBy and TopByDescending
These may sound like they belong in the FooBy section, but theyre somewhat different: theyre effectively specializations of OrderBy and OrderByDescending where you already know how many elements you want to preserve. The return type would be IOrderedEnumerable<T> so you could still use ThenBy/ThenByDescending as normal. That would make the following two queries equivalent but the second _might_ be a lot more efficient than the first:
var takeQuery = people.OrderBy(p => p.LastName)
                      .ThenBy(p => p.FirstName)
                      .Take(3);
var topQuery = people.TopBy(p => p.LastName, 3)
                     .ThenBy(p => p.FirstName);
An implementation could easily delegate to various different strategies depending on the number given for example, if you asked for more than 10 values, it may not be worth doing anything more than a simple sort and restrict the output. If you asked for just the top 3 values, that could return an IOrderedEnumerable implementation specifically hard-coded to 3 values, etc.
Aside from anything else, if you were confident in what the implementation did (and thats a _very_ big "if") you could use a potentially huge input sequence with such a query larger than you could fit into memory in one go. Thats fine if youre only keeping the top three values youve seen so far, but would fail for a complete ordering, even one which was able to yield results before performing _all_ the ordering: if it doesnt _know_ youre going to stop after three elements, it cant throw anything away.
Perhaps this is too specialized an operator but its an interesting one to think about. Its worth noting that this probably only makes sense for LINQ to Objects, which never gets to see the whole query in one go. Providers like LINQ to SQL can optimize queries of the form OrderBy(…).ThenBy(…).Take(…) because by the time they need to translate the query into SQL, they will have an expression tree representation which includes the "Take" part.
### TryFastCount and TryFastElementAt
One of the implementation details of Edulinq is its TryFastCount method, which basically encapsulates the logic around attempting to find the count of a sequence if it implements ICollection or ICollection<T>. Various built-in LINQ operators find this useful, and anyone writing their own operators has a reasonable chance of bumping into it as well. It seems pointless to duplicate the code all over the place… why not expose it? The signatures might look something like this:
public static bool TryFastCount<TSource>(
    this IEnumerable<TSource> source,
    out int count)
public static bool TryFastElementAt<TSource>(
    this IEnumerable<TSource> source,
    int index,
    out TSource value)
I would expect TryFastElementAt to use the indexer if the sequence implemented IList<T> without performing any validation: that ought to be the responsibility of the caller. TryFastCount could use a Nullable<int> return type instead of the return value / out parameter split, but Ive kept it consistent with the methods which exist elsewhere in the framework
### Scan and SelectAdjacent
These are related operators in that they deal with wanting a more global view than just the current element. Scan would act similarly to Aggregate except that it would yield the accumulator value after each element. Heres an example of keeping a running total:
// Signature:
public static IEnumerable<TAccumulate> Scan<TSource, TAccumulate>(
    this IEnumerable<TSource> source,
    TAccumulate seed,
    Func<TAccumulate, TSource, TAccumulate> func)
    
int[] source = new int[] { 3, 5, 2, 1, 4 };
var query = source.Scan(0, (current, item) => current + item);
query.AssertSequenceEqual(3, 8, 10, 11, 15);
There _could_ be a more complicated overload with an extra conversion from TAccumulate to an extra TResult type parameter. That would let us write a Fibonacci sequence query in one line, if we really wanted to…
The SelectAdjacent operator would simply present a selector function with pairs of adjacent items. Heres a similar example, this time calculating the difference between each pair:
// Signature:
public static IEnumerable<TResult> SelectAdjacent<TSource, TResult>(
    this IEnumerable<TSource> source,
    Func<TSource, TSource, TResult> selector)
    
int[] source = new int[] { 3, 5, 2, 1, 4 };
var query = source.SelectAdjacent((current, next) => next current);
query.AssertSequenceEqual(2, -3, -1, 3);
One oddity here is that the result sequence always contains one item fewer than the source sequence. If we wanted to keep the length the same, there are various approaches we could take but the best one would depend on the situation.
This sounds like a pretty obscure operator, but Ive actually seen quite a few LINQ questions on Stack Overflow where it could have been useful. Is it useful often _enough_ to deserve its own operator? Maybe… maybe not.
### DelimitWith
This one is really just a bit of a peeve but again, its a pretty common requirement. We often want to take a sequence and create a single string which is (say) a comma-delimited version. Yay, String.Join does exactly what we need particularly in .NET 4, where theres an [overload taking IEnumerable<T>](http://msdn.microsoft.com/en-us/library/dd992421.aspx) so you dont need to convert it to a string array first. However, its still a _static_ method on string and the name "Join" also looks slightly odd in the context of a LINQ query, as its got nothing to do with a LINQ-style join.
Compare these two queries: which do you think reads better, and feels more "natural" in LINQ?
// Current state of play…
var names = string.Join(",",
                        people.Where(p => p.Age < 18)
                              .Select(p => p.FirstName));
// Using DelimitWith
var names = people.Where(p => p.Age < 18)
                  .Select(p => p.FirstName)
                  .DelimitWith(",");
I know which I prefer :)
### ToHashSet
(Added on February 23rd 2011.)
Im surprised I missed this one first time round Ive bemoaned its omission in various places before now. Its easy to create a list, dictionary, lookup or array from an anonymous type, but you cant create a set that way. Thats mad, given how simple the relevant operator is, even with an overload for a custom equality comparer:
public static HashSet<TSource> ToHashSet<TSource>(
    this IEnumerable<TSource> source)
{
    return source.ToHashSet(EqualityComparer<TSource>.Default);
}
public static HashSet<TSource> ToHashSet<TSource>(
    this IEnumerable<TSource> source,
    IEqualityComparer<TSource> comparer)
{
    if (source == null)
    {
        throw new ArgumentNullException("source");
    }
    return new HashSet<TSource>(source, comparer ?? EqualityComparer<TSource>.Default);
}
This also makes it much simpler to create a HashSet in a readable way from an existing query expression, without either wrapping the whole query in the constructor call or using a local variable.
### Conclusion
These are just the most useful extra methods I thought of, based on the kinds of query folks on Stack Overflow have asked about. I think its interesting that some are quite general MaxBy, ExceptBy, Scan and so on whereas others (TopBy, SelectAdjacent and particularly DelimitWith) are simply aimed at making some very specific but common situations simpler. It feels to me like the more general operators really are _missing_ from LINQ they would fit quite naturally but the more specific ones probably deserve to be in a separate static class, as "extras".
This is only scratching the surface of whats possible, of course System.Interactive.EnumerableEx in Reactive Extensions has _loads_ of options. Some of them are deliberate parallels of the operators in Observable, but plenty make sense on their own too.
One operator you may have _expected_ to see in this list is ForEach. This is a controversial topic, but Eric Lippert [has written about it very clearly](http://blogs.msdn.com/b/ericlippert/archive/2009/05/18/foreach-vs-foreach.aspx) (no surprise there, then). Fundamentally LINQ is about _querying_ a sequence, not taking _action_ on it. ForEach breaks that philosophy, which is why I havent included it here. Usually a foreach statement is a perfectly good alternative, and make the "action" aspect clearer.
[C#](https://codeblog.jonskeet.uk/category/csharp/), [Edulinq](https://codeblog.jonskeet.uk/category/edulinq/), [LINQ](https://codeblog.jonskeet.uk/category/linq/)
# [Reimplementing LINQ to Objects: Part 37 Guiding principles](https://codeblog.jonskeet.uk/2011/01/20/reimplementing-linq-to-objects-part-37-guiding-principles/)
[January 20, 2011](https://codeblog.jonskeet.uk/2011/01/20/reimplementing-linq-to-objects-part-37-guiding-principles/) [jonskeet](https://codeblog.jonskeet.uk/author/jonskeet/) [5 Comments](https://codeblog.jonskeet.uk/2011/01/20/reimplementing-linq-to-objects-part-37-guiding-principles/#comments)
Now that Im "done" reimplementing LINQ to Objects in that Ive implemented all the methods in System.Linq.Enumerable I wanted to write a few posts looking at the bigger picture. Im not 100% sure of what this will consist of yet; I want to avoid this blog series continuing forever. However, Im confident it will contain (in no particular order):
- This post: principles governing the behaviour of LINQ to Objects
- Missing operators: what else Id have liked to see in Enumerable
- Optimization: where the .NET implementation could be further optimized, and why some obvious-sounding optimizations may be inappropriate
- How query expression translations work, in brief (and with a cheat sheet)
- The difference between IQueryable<T> and IEnumerable<T>
- Sequence identity, the "Contains" issue, and other knotty design questions
- Running the Edulinq tests against other implementations
If there are other areas you want me to cover, please let me know.
### The principles behind the LINQ to Objects implementation
The design LINQ to Objects is built on a few guiding principles, both in terms of design and implementation details. You need to understand these, but also implementations should be clear about what theyre doing in these terms too.
### Extension method targets and argument validation
IEnumerable<T> is the core sequence type, not just for LINQ but for .NET as a whole. Almost _everything_ is written in terms of IEnumerable<T> at least as input, with the following exceptions:
- Empty, Range and Repeat dont have input sequences (these are the only non-extension methods)
- OfType and Cast work on the non-generic IEnumerable type instead
- ThenBy and ThenByDescending work on IOrderedEnumerable<T>
All operators other than AsEnumerable verify that any input sequence is non-null. This validation is performed eagerly (i.e. when the method is called) even if the operator uses deferred execution for the results. Any delegate used (typically a projection or predicate of some kind) must be non-null. Again, this validation is performed eagerly.
IEqualityComparer<T> is used for all custom equality comparisons. Any parameter of this type _may_ be null, in which case the default equality comparer for the type is used. In _most_ cases the default equality comparer for the type is also used when no custom equality comparer is used, but Contains has some odd behaviour around this. Equality comparers are expected to be able to handle null values. IComparer<T> is _only_ used by the OrderBy/ThenBy operators and their descending counterparts and only then if you want custom comparisons between keys. Again, a null IComparer<T> means "use the default for the type"
### Timing of input sequence "opening"
Any operator with a return type of IEnumerable<T> or IOrderedEnumerable<T> uses _deferred execution_. This means that the method doesnt read anything from any input sequences until someone starts reading from the result sequence. Its not clearly defined exactly _when_ input sequences will first be accessed for some operators if may be when GetEnumerator() is called; for others it may be on the first call to MoveNext() on the resulting iterator. Callers should not depend on these slight variations. Deferred execution is common for operators in the middle of queries. Operators which use deferred execution effectively represent queries rather than the results of queries so if you change the contents of the original source of the query and then iterate over the query itself again, youll see the change. For example:
List<string> source = new List<string>();
var query = source.Select(x => x.ToUpper());
        
// This loop wont write anything out
foreach (string x in query)
{
    Console.WriteLine(x);
}
        
source.Add("foo");
source.Add("bar");
// This loop will write out "FOO" and "BAR" even
// though we havent changed the value of "query"
foreach (string x in query)
{
    Console.WriteLine(x);
}
Deferred execution is one of the hardest parts of LINQ to understand, but once you do, everything becomes somewhat simpler.
All other operators use _immediate execution_, fetching all the data they need from the input before they return a value… so that by the time they _do_ return, they will no longer see or care about changes to the input sequence. For operators returning a scalar value (such as Sum and Average) this is blatantly obvious the value of a variable of type double isnt going to change just because youve added something to a list. However, its slightly less for the "ToXXX" methods: ToLookup, ToArray, ToList and ToDictionary. These do _not_ return views on the original sequence, unlike the "As" methods: AsEnumerable which weve seen, and Queryable.AsQueryable which I didnt implement. Focus on the prefix part of the name: the "To" part indicates a conversion to a particular type. The "As" prefix indicates a wrapper of some kind. This is consistent with other parts of the framework, such as List<T>.AsReadOnly and Array.AsReadOnly<T>.
Very importantly, LINQ to Objects only iterates over any input sequence at most **once**, whether the execution is deferred or immediate. Some operators would be easier to implement if you could iterate over the input twice but its important that they dont do so. Of course if you provide the same sequence for two inputs, it will treat those as logically different sequences. Similarly if you iterate over a result sequence more than once (for operators that return IEnumerable<T> or a related interface, rather than List<T> or an array etc), that will iterate over the input sequence again.
This means its fine to use LINQ to Objects with sequences which may only be read once (such as a network stream), or which are relatively expensive to reread (imagine a log file reader over a huge set of logs) or which give inconsistent results (imagine a sequence of random numbers). In some cases its okay to use LINQ to Objects with an infinite sequence in others its not. Its _usually_ fairly obvious which is the case.
### Timing of input sequence reading, and memory usage
Where possible within deferred execution, operators act in a _streaming_ fashion, only reading from the input sequence when they have to, and "forgetting" data as soon as they can. This allows for long potentially infinite sequences to be handled elegantly without memory running out.
Some operators naturally need to read all the data in before they can return anything. The most obvious example of this is Reverse, which will always yield the last element of the input stream as the first element in the result stream.
A third pattern occurs with operators such as Distinct, which yield data as they go, but accumulate elements too, taking more and more memory until the caller stops iterating (usually either by jumping out of the foreach loop, or letting it terminate naturally).
Where an operator takes two input sequences such as Join you need to understand the consumption of each one separately. For example, Join uses deferred execution, but as soon as you ask for the first element of the result set, it will read the "second" sequence _completely_ and buffer it whereas the "first" sequence is streamed. This isnt the case for all operators with two inputs, of course Zip streams both input sequences, for example. Check the documentation and the relevant Edulinq blog post for details.
Obviously any operator which uses immediate execution has to read all the data its interested in before it returns. This doesnt necessarily mean they will read to the end of the sequence though, and they may not need to buffer the data they read. (Simple examples are ToList which has to keep everything, and Sum which doesnt.)
### Queries vs data
Closely related to the details of when the input is read is the concept of what the result of an operator actually represents. Operators which use deferred execution return _queries_: each time you iterate over the result sequence, the query will look at the input sequence again. The query itself doesnt contain the data it just knows how to get at the data.
Operators which use immediate execution work the other way round: they read all the data they need, and then forget about the input sequence. For operators like Average and Sum this is obvious as its just a simple scalar value but for operators like ToList, ToDictionary, ToLookup and ToArray, it means that the operator has to make a copy of everything it needs. (This is potentially a _shallow_ copy of course depending on what user-defined projections are applied. The normal behaviour of mutable reference types is still valid.)
I realise that in many ways Ive just said the same thing multiple times now but hopefully that will help this crucial aspect of LINQ behaviour sink in, if you were still in any doubt.
### Exception handling
Im unaware of any situation in which LINQ to Objects will catch an exception. If your predicate or projection throws an exception, it will propagate in the obvious way.
However, LINQ to Objects _does_ ensure that any iterator it reads from is disposed appropriately assuming that the caller disposes of any result sequences properly, of course. Note that the foreach statement implicitly disposes of the iterator in a finally block.
### Optimization
Various operators are optimized when they detect at execution time that the input sequence theyre working on offers a shortcut.
The types most commonly detected are:
- ICollection<T> and ICollection for their Count property
- IList<T> for its random access indexer
Ill look at optimization in much more detail in a separate post.
### Conclusion
This post has not been around the guiding principles behind LINQ itself lambda calculus or anything like that. Its more been a summary of the various aspects of behaviour weve seen across the various operators weve implemented. Theyre the rules Ive had to follow in order to make Edulinq reasonably consistent with LINQ to Objects.
Next time Ill talk about some of the operators which I think _should_ have made it into the core framework, at least for LINQ to Objects.

View File

@@ -0,0 +1,384 @@
Clipped from: [https://devblogs.microsoft.com/dotnet/how-async-await-really-works/](https://devblogs.microsoft.com/dotnet/how-async-await-really-works/)
![Exported image](Exported%20image%2020240808113926-0.jpeg)
Stephen Toub - MSFT
Several weeks ago, the [.NET Blog](https://devblogs.microsoft.com/dotnet/) featured a post [What is .NET, and why should you choose it?](https://devblogs.microsoft.com/dotnet/why-dotnet/). It provided a high-level overview of the platform, summarizing various components and design decisions, and promising more in-depth posts on the covered areas. This post is the first such follow-up, deep-diving into the history leading to, the design decisions behind, and implementation details of async/await in C# and .NET.
The support for async/await has been around now for over a decade. In that time, its transformed how scalable code is written for .NET, and its both viable and extremely common to utilize the functionality without understanding exactly whats going on under the covers. You start with a synchronous method like the following (this method is “synchronous” because a caller will not be able to do anything else until this whole operation completes and control is returned back to the caller):
// Synchronously copy all data from source to destination.public void CopyStreamToStream(Stream source, Stream destination){ var buffer = new byte[0x1000]; int numRead; while ((numRead = source.Read(buffer, 0, buffer.Length)) != 0) { destination.Write(buffer, 0, numRead); }}
Then you sprinkle a few keywords, change a few method names, and you end up with the following asynchronous method instead (this method is “asynchronous” because control is expected to be returned back to its caller very quickly and possibly before the work associated with the whole operation has completed):
// Asynchronously copy all data from source to destination.public async Task CopyStreamToStreamAsync(Stream source, Stream destination){ var buffer = new byte[0x1000]; int numRead; while ((numRead = await source.ReadAsync(buffer, 0, buffer.Length)) != 0) { await destination.WriteAsync(buffer, 0, numRead); }}
Almost identical in syntax, still able to utilize all of the same control flow constructs, but now non-blocking in nature, with a significantly different underlying execution model, and with all the heavy lifting done for you under the covers by the C# compiler and core libraries.
While its common to use this support without knowing exactly whats happening under the hood, Im a firm believer that understanding how something actually works helps you to make even better use of it. For async/await in particular, understanding the mechanisms involved is especially helpful when you want to look below the surface, such as when youre trying to debug things gone wrong or improve the performance of things otherwise gone right. In this post, then, well deep-dive into exactly how await works at the language, compiler, and library level, so that you can make the most of these valuable features.
To do that well, though, we need to go way back to before async/await to understand what state-of-the-art asynchronous code looked like in its absence. Fair warning, it wasnt pretty.
## In the beginning…
All the way back in .NET Framework 1.0, there was the Asynchronous Programming Model pattern, otherwise known as the APM pattern, otherwise known as the Begin/End pattern, otherwise known as the IAsyncResult pattern. At a high-level, the pattern is simple. For a synchronous operation DoStuff:
class Handler{ public int DoStuff(string arg);}
there would be two corresponding methods as part of the pattern: a BeginDoStuff method and an EndDoStuff method:
class Handler{ public int DoStuff(string arg); public IAsyncResult BeginDoStuff(string arg, AsyncCallback? callback, object? state); public int EndDoStuff(IAsyncResult asyncResult);}
BeginDoStuff would accept all of the same parameters as does DoStuff, but in addition it would also accept an [AsyncCallback](https://github.com/dotnet/runtime/blob/967a59712996c2cdb8ce2f65fb3167afbd8b01f3/src/libraries/System.Private.CoreLib/src/System/AsyncCallback.cs#L14) delegate and an opaque state object, one or both of which could be null. The Begin method was responsible for initiating the asynchronous operation, and if provided with a callback (often referred to as the “continuation” for the initial operation), it was also responsible for ensuring the callback was invoked when the asynchronous operation completed. The Begin method would also construct an instance of a type that implemented [IAsyncResult](https://github.com/dotnet/runtime/blob/967a59712996c2cdb8ce2f65fb3167afbd8b01f3/src/libraries/System.Private.CoreLib/src/System/IAsyncResult.cs#L17-L27), using the optional state to populate that IAsyncResults AsyncState property:
namespace System{ public interface IAsyncResult { object? AsyncState { get; } WaitHandle AsyncWaitHandle { get; } bool IsCompleted { get; } bool CompletedSynchronously { get; } } public delegate void AsyncCallback(IAsyncResult ar);}
This IAsyncResult instance would then both be returned from the Begin method as well as passed to the AsyncCallback when it was eventually invoked. When ready to consume the results of the operation, a caller would then pass that IAsyncResult instance to the End method, which was responsible for ensuring the operation was completed (synchronously waiting for it to complete by blocking if it wasnt) and then returning any result of the operation, including propagating any errors/exceptions that may have occurred. Thus, instead of writing code like the following to perform the operation synchronously:
try{ int i = handler.DoStuff(arg); Use(i);}catch (Exception e){ ... // handle exceptions from DoStuff and Use}
the Begin/End methods could be used in the following manner to perform the same operation asynchronously:
try{ handler.BeginDoStuff(arg, iar => { try { Handler handler = (Handler)iar.AsyncState!; int i = handler.EndDoStuff(iar); Use(i); } catch (Exception e2) { ... // handle exceptions from EndDoStuff and Use } }, handler);}catch (Exception e){ ... // handle exceptions thrown from the synchronous call to BeginDoStuff}
For anyone whos dealt with callback-based APIs in any language, this should feel familiar.
Things only got more complicated from there, however. For instance, theres the issue of “stack dives.” A stack dive is when code repeatedly makes calls that go deeper and deeper on the stack, to the point where it could potentially stack overflow. The Begin method is allowed to invoke the callback synchronously if the operation completes synchronously, meaning the call to Begin might itself directly invoke the callback. And “asynchronous” operations that complete synchronously are actually very common; theyre not “asynchronous” because theyre guaranteed to complete asynchronously but rather are just permitted to. For example, consider an asynchronous read from some networked operation, like receiving from a socket. If you need only a small amount of data for each individual operation, such as reading some header data from a response, you might put a buffer in place in order to avoid the overhead of lots of system calls. Instead of doing a small read for just the amount of data you need immediately, you perform a larger read into the buffer and then consume data from that buffer until its exhausted; that lets you reduce the number of expensive system calls required to actually interact with the socket. Such a buffer might exist behind whatever asynchronous abstraction youre using, such that the first “asynchronous” operation you perform (filling the buffer) completes asynchronously, but then all subsequent operations until that underlying buffer is exhausted dont actually need to do any I/O, instead just pulling from the buffer, and can thus all complete synchronously. When the Begin method performs one of these operations, and finds it completes synchronously, it can then invoke the callback synchronously. That means you have one stack frame that called the Begin method, another stack frame for the Begin method itself, and now another stack frame for the callback. Now what happens if that callback turns around and calls Begin again? If that operation completes synchronously and its callback is invoked synchronously, youre now again several more frames deep on the stack. And so on, and so on, until eventually you run out of stack.
This is a real possibility thats easy to repro. Try this program on .NET Core:
using System.Net;using System.Net.Sockets;using Socket listener = new Socket(AddressFamily.InterNetwork, SocketType.Stream, ProtocolType.Tcp);listener.Bind(new IPEndPoint(IPAddress.Loopback, 0));listener.Listen();using Socket client = new Socket(AddressFamily.InterNetwork, SocketType.Stream, ProtocolType.Tcp);client.Connect(listener.LocalEndPoint!);using Socket server = listener.Accept();_ = server.SendAsync(new byte[100_000]);var mres = new ManualResetEventSlim();byte[] buffer = new byte[1];var stream = new NetworkStream(client);void ReadAgain(){ stream.BeginRead(buffer, 0, 1, iar => { if (stream.EndRead(iar) != 0) { ReadAgain(); // uh oh! } else { mres.Set(); } }, null);};ReadAgain();mres.Wait();
Here Ive set up a simple client socket and server socket connected to each other. The server sends 100,000 bytes to the client, which then proceeds to use BeginRead/EndRead to consume them “asynchronously” one at a time (this is terribly inefficient and is only being done in the name of pedagogy). The callback passed to BeginRead finishes the read by calling EndRead, and then if it successfully read the desired byte (in which case it wasnt yet at end-of-stream), it issues another BeginRead via a recursive call to the ReadAgain local function. However, in .NET Core, socket operations are much faster than they were on .NET Framework, and will complete synchronously if the OS is able to satisfy the operation synchronously (noting the kernel itself has a buffer used to satisfy socket receive operations). Thus, this stack overflows:
[![Stack overflow due to improper handling of synchronous completion](Exported%20image%2020240808113926-1.png)](https://devblogs.microsoft.com/dotnet/wp-content/uploads/sites/10/2023/03/BeginReadStackOverflow.png)
So, compensation for this was built into the APM model. There are two possible ways to compensate for this:
1. Dont allow the AsyncCallback to be invoked synchronously. If its always invoked asynchronously, even if the operation completes synchronously, then the risk of stack dives goes away. But so too does performance, because operations that complete synchronously (or so quickly that theyre observably indistinguishable) are very common, and forcing each of those to queue its callback adds measurable overhead.
2. Employ a mechanism that allows the caller rather than the callback to do the continuation work if the operation completes synchronously. That way, you escape the extra method frame and continue doing the follow-on work no deeper on the stack.
The APM pattern goes with option (2). For that, the IAsyncResult interface exposes two related but distinct members: IsCompleted and CompletedSynchronously. IsCompleted tells you whether the operation has completed: you can check it multiple times, and eventually itll transition from false to true and then stay there. In contrast, CompletedSynchronously never changes (or if it does, its a nasty bug waiting to happen); its used to communicate between the caller of the Begin method and the AsyncCallback which of them is responsible for performing any continuation work. If CompletedSynchronously is false, then the operation is completing asynchronously and any continuation work in response to the operation completing should be left up to the callback; after all, if the work didnt complete synchronously, the caller of Begin cant really handle it because the operation isnt known to be done yet (and if the caller were to just call End, it would block until the operation completed). If, however, CompletedSynchronously is true, if the callback were to handle the continuation work, then it risks a stack dive, as itll be performing that continuation work deeper on the stack than where it started. Thus, any implementations at all concerned about such stack dives need to examine CompletedSynchronously and have the caller of the Begin method do the continuation work if its true, which means the callback then needs to _not_ do the continuation work. This is also why CompletedSynchronously must never change: the caller and the callback need to see the same value to ensure that the continuation work is performed once and only once, regardless of race conditions.
In our previous DoStuff example, that then leads to code like this:
try{ IAsyncResult ar = handler.BeginDoStuff(arg, iar => { if (!iar.CompletedSynchronously) { try { Handler handler = (Handler)iar.AsyncState!; int i = handler.EndDoStuff(iar); Use(i); } catch (Exception e2) { ... // handle exceptions from EndDoStuff and Use } } }, handler); if (ar.CompletedSynchronously) { int i = handler.EndDoStuff(ar); Use(i); }}catch (Exception e){ ... // handle exceptions that emerge synchronously from BeginDoStuff and possibly EndDoStuff/Use}
Thats a mouthful. And so far weve only looked at consuming the pattern… we havent looked at implementing the pattern. While most developers wouldnt need to be concerned about leaf operations (e.g. implementing the actual Socket.BeginReceive/EndReceive methods that interact with the operating system), many, many developers would need to be concerned with composing these operations (performing multiple asynchronous operations that together form a larger one), which means not only consuming other Begin/End methods but also implementing them yourself so that your composition itself can be consumed elsewhere. And, youll notice there was no control flow in my previous DoStuff example. Introduce multiple operations into this, especially with even simple control flow like a loop, and all of a sudden this becomes the domain of experts that enjoy pain, or blog post authors trying to make a point.
So just to drive that point home, lets implement a complete example. At the beginning of this post, I showed a CopyStreamToStream method that copies all of the data from one stream to another (à la Stream.CopyTo, but, for the sake of explanation, assuming that doesnt exist):
public void CopyStreamToStream(Stream source, Stream destination){ var buffer = new byte[0x1000]; int numRead; while ((numRead = source.Read(buffer, 0, buffer.Length)) != 0) { destination.Write(buffer, 0, numRead); }}
Straightforward: we repeatedly read from one stream and then write the resulting data to the other, read from one stream and write to the other, and so on, until we have no more data to read. Now, how would we implement this asynchronously using the APM pattern? Something like this:
public IAsyncResult BeginCopyStreamToStream( Stream source, Stream destination, AsyncCallback callback, object state){ var ar = new MyAsyncResult(state); var buffer = new byte[0x1000]; Action<IAsyncResult?> readWriteLoop = null!; readWriteLoop = iar => { try { for (bool isRead = iar == null; ; isRead = !isRead) { if (isRead) { iar = source.BeginRead(buffer, 0, buffer.Length, static readResult => { if (!readResult.CompletedSynchronously) { ((Action<IAsyncResult?>)readResult.AsyncState!)(readResult); } }, readWriteLoop); if (!iar.CompletedSynchronously) { return; } } else { int numRead = source.EndRead(iar!); if (numRead == 0) { ar.Complete(null); callback?.Invoke(ar); return; } iar = destination.BeginWrite(buffer, 0, numRead, writeResult => { if (!writeResult.CompletedSynchronously) { try { destination.EndWrite(writeResult); readWriteLoop(null); } catch (Exception e2) { ar.Complete(e); callback?.Invoke(ar); } } }, null); if (!iar.CompletedSynchronously) { return; } destination.EndWrite(iar); } } } catch (Exception e) { ar.Complete(e); callback?.Invoke(ar); } }; readWriteLoop(null); return ar;}public void EndCopyStreamToStream(IAsyncResult asyncResult){ if (asyncResult is not MyAsyncResult ar) { throw new ArgumentException(null, nameof(asyncResult)); } ar.Wait();}private sealed class MyAsyncResult : IAsyncResult{ private bool _completed; private int _completedSynchronously; private ManualResetEvent? _event; private Exception? _error; public MyAsyncResult(object? state) => AsyncState = state; public object? AsyncState { get; } public void Complete(Exception? error) { lock (this) { _completed = true; _error = error; _event?.Set(); } } public void Wait() { WaitHandle? h = null; lock (this) { if (_completed) { if (_error is not null) { throw _error; } return; } h = _event ??= new ManualResetEvent(false); } h.WaitOne(); if (_error is not null) { throw _error; } } public WaitHandle AsyncWaitHandle { get { lock (this) { return _event ??= new ManualResetEvent(_completed); } } } public bool CompletedSynchronously { get { lock (this) { if (_completedSynchronously == 0) { _completedSynchronously = _completed ? 1 : -1; } return _completedSynchronously == 1; } } } public bool IsCompleted { get { lock (this) { return _completed; } } }}
Yowsers. And, even with all of that gobbledygook, its still not a great implementation. For example, the IAsyncResult implementation is locking on every operation rather than doing things in a more lock-free manner where possible, the Exception is being stored raw rather than as an [ExceptionDispatchInfo](https://github.com/dotnet/runtime/blob/967a59712996c2cdb8ce2f65fb3167afbd8b01f3/src/libraries/System.Private.CoreLib/src/System/Runtime/ExceptionServices/ExceptionDispatchInfo.cs#L9-L16) that would enable augmenting its call stack when propagated, theres a lot of allocation involved in each individual operation (e.g. a delegate being allocated for each BeginWrite call), and so on. Now, imagine having to do all of this for each method you wanted to write. Every time you wanted to write a reusable method that would consume another asynchronous operation, youd need to do all of this work. And if you wanted to write reusable combinators that could operate over multiple discrete IAsyncResults efficiently (think Task.WhenAll), thats another level of difficulty; every operation implementing and exposing its own APIs specific to that operation meant there was no lingua franca for talking about them all similarly (though some developers wrote libraries that tried to ease the burden a bit, typically via another layer of callbacks that enabled the API to supply an appropriate AsyncCallback to a Begin method).
And all of that complication meant that very few folks even attempted this, and for those who did, well, bugs were rampant. To be fair, this isnt really a criticism of the APM pattern. Rather, its a critique of callback-based asynchrony in general. Were all so used to the power and simplicity that control flow constructs in modern languages provide us with, and callback-based approaches typically run afoul of such constructs once any reasonable amount of complexity is introduced. No other mainstream language had a better alternative available, either.
We needed a better way, one in which we learned from the APM pattern, incorporating the things it got right while avoiding its pitfalls. An interesting thing to note is that the APM pattern is just that, a pattern; the runtime, core libraries, and compiler didnt provide any assistance in consuming or implementing the pattern.
## Event-Based Asynchronous Pattern
.NET Framework 2.0 saw a few APIs introduced that implemented a different pattern for handling asynchronous operations, one primarily intended for doing so in the context of client applications. This Event-based Asynchronous Pattern, or EAP, also came as a pair of members (at least, possibly more), this time a method to initiate the asynchronous operation and an event to listen for its completion. Thus, our earlier DoStuff example might have been exposed as a set of members like this:
class Handler{ public int DoStuff(string arg); public void DoStuffAsync(string arg, object? userToken); public event DoStuffEventHandler? DoStuffCompleted;}public delegate void DoStuffEventHandler(object sender, DoStuffEventArgs e);public class DoStuffEventArgs : AsyncCompletedEventArgs{ public DoStuffEventArgs(int result, Exception? error, bool canceled, object? userToken) : base(error, canceled, usertoken) => Result = result; public int Result { get; }}
Youd register your continuation work with the DoStuffCompleted event and then invoke the DoStuffAsync method; it would initiate the operation, and upon that operations completion, the DoStuffCompleted event would be raised asynchronously from the caller. The handler could then run its continuation work, likely validating that the userToken supplied matched the one it was expecting, enabling multiple handlers to be hooked up to the event at the same time.
This pattern made a few use cases a bit easier while making other uses cases significantly harder (and given the previous APM CopyStreamToStream example, thats saying something). It didnt get rolled out in a widespread manner, and it came and went effectively in a single release of .NET Framework, albeit leaving behind the APIs added during its tenure, like Ping.SendAsync/Ping.PingCompleted:
public class Ping : Component{ public void SendAsync(string hostNameOrAddress, object? userToken); public event PingCompletedEventHandler? PingCompleted; ...}
However, it did add one notable advance that the APM pattern didnt factor in at all, and that has endured into the models we embrace today: [SynchronizationContext](https://github.com/dotnet/runtime/blob/967a59712996c2cdb8ce2f65fb3167afbd8b01f3/src/libraries/System.Private.CoreLib/src/System/Threading/SynchronizationContext.cs#L6).
SynchronizationContext was also introduced in .NET Framework 2.0, as an abstraction for a general scheduler. In particular, SynchronizationContexts most used method is Post, which queues a work item to whatever scheduler is represented by that context. The base implementation of SynchronizationContext, for example, just represents the ThreadPool, and so the [base implementation of](https://github.com/dotnet/runtime/blob/95df571be36ed8973d09746b61fae16b2e3f251f/src/libraries/System.Private.CoreLib/src/System/Threading/SynchronizationContext.cs#L22) SynchronizationContext.Post simply delegates to [ThreadPool.QueueUserWorkItem](https://learn.microsoft.com/dotnet/api/system.threading.threadpool.queueuserworkitem), which is used to ask the ThreadPool to invoke the supplied callback with the associated state on one the pools threads. However, SynchronizationContexts bread-and-butter isnt just about supporting arbitrary schedulers, rather its about supporting scheduling in a manner that works according to the needs of various application models.
Consider a UI framework like Windows Forms. As with most UI frameworks on Windows, controls are associated with a particular thread, and that thread runs a message pump which runs work thats able to interact with those controls: only that thread should try to manipulate those controls, and any other thread that wants to interact with the controls should do so by sending a message to be consumed by the UI threads pump. Windows Forms makes this easy with methods like Control.BeginInvoke, which queues the supplied delegate and arguments to be run by whatever thread is associated with that Control. You can thus write code like this:
private void button1_Click(object sender, EventArgs e){ ThreadPool.QueueUserWorkItem(_ => { string message = ComputeMessage(); button1.BeginInvoke(() => { button1.Text = message; }); });}
That will offload the ComputeMessage() work to be done on a ThreadPool thread (so as to keep the UI responsive while its being processed), and then when that work has completed, queue a delegate back to the thread associated with button1 to update button1s label. Easy enough. WPF has something similar, just with its Dispatcher type:
private void button1_Click(object sender, RoutedEventArgs e){ ThreadPool.QueueUserWorkItem(_ => { string message = ComputeMessage(); button1.Dispatcher.InvokeAsync(() => { button1.Content = message; }); });}
And .NET MAUI has something similar. But what if I wanted to put this logic into a helper method? e.g.
// Call ComputeMessage and then invoke the update action to update controls.internal static void ComputeMessageAndInvokeUpdate(Action<string> update) { ... }
I could then use that like this:
private void button1_Click(object sender, EventArgs e){ ComputeMessageAndInvokeUpdate(message => button1.Text = message);}
but how could ComputeMessageAndInvokeUpdate be implemented in such a way that it could work in any of those applications? Would it need to be hardcoded to know about every possible UI framework? Thats where SynchronizationContext shines. We might implement the method like this:
internal static void ComputeMessageAndInvokeUpdate(Action<string> update){ SynchronizationContext? sc = SynchronizationContext.Current; ThreadPool.QueueUserWorkItem(_ => { string message = ComputeMessage(); if (sc is not null) { sc.Post(_ => update(message), null); } else { update(message); } });}
That uses the SynchronizationContext as an abstraction to target whatever “scheduler” should be used to get back to the necessary environment for interacting with the UI. Each application model then ensures its published as SynchronizationContext.Current a SynchronizationContext-derived type that does the “right thing.” For example, [Windows Forms has this](https://github.com/dotnet/winforms/blob/41b11b6a7290a2bbc0c293042f30d9632e55aae2/src/System.Windows.Forms/src/System/Windows/Forms/WindowsFormsSynchronizationContext.cs#L13):
public sealed class WindowsFormsSynchronizationContext : SynchronizationContext, IDisposable{ public override void Post(SendOrPostCallback d, object? state) => _controlToSendTo?.BeginInvoke(d, new object?[] { state }); ...}
and [WPF has this](https://github.com/dotnet/wpf/blob/c67b9f6f5ad04f5c264b52de0733a8832714615f/src/Microsoft.DotNet.Wpf/src/WindowsBase/System/Windows/Threading/DispatcherSynchronizationContext.cs#L18):
public sealed class DispatcherSynchronizationContext : SynchronizationContext{ public override void Post(SendOrPostCallback d, Object state) => _dispatcher.BeginInvoke(_priority, d, state); ...}
ASP.NET _used_ to [have one](https://referencesource.microsoft.com/#System.Web/AspNetSynchronizationContext.cs,16), which didnt actually care about what thread work ran on, but rather that work associated with a given request was serialized such that multiple threads wouldnt concurrently be accessing a given HttpContext:
internal sealed class AspNetSynchronizationContext : AspNetSynchronizationContextBase{ public override void Post(SendOrPostCallback callback, Object state) => _state.Helper.QueueAsynchronous(() => callback(state)); ...}
This also isnt limited to such main application models. For example, [xunit](https://github.com/xunit/xunit) is a popular unit testing framework, one that .NETs core repos use for their unit testing, and it also employs multiple custom SynchronizationContexts. You can, for example, allow tests to run in parallel but limit the number of tests that are allowed to be running concurrently. How is that enabled? Via a SynchronizationContext:
public class MaxConcurrencySyncContext : SynchronizationContext, IDisposable{ public override void Post(SendOrPostCallback d, object? state) { var context = ExecutionContext.Capture(); workQueue.Enqueue((d, state, context)); workReady.Set(); }}
[MaxConcurrencySyncContext](https://github.com/xunit/xunit/blob/601e2d830853fa2ef0048d34afae520d6b73deca/src/xunit.v3.core/Sdk/MaxConcurrencySyncContext.cs#L14)s Post method just queues the work to its own internal work queue, which it then processes on its own worker threads, where it controls how many there are based on the max concurrency desired. You get the idea.
How does this tie in with the Event-based Asynchronous Pattern? Both EAP and SynchronizationContext were introduced at the same time, and the EAP dictated that the completion events should be queued to whatever SynchronizationContext was current when the asynchronous operation was initiated. To simplify that ever so slightly (and arguably not enough to warrant the extra complexity), some helper types were also introduced in System.ComponentModel, in particular AsyncOperation and AsyncOperationManager. The former was just a tuple that wrapped the user-supplied state object and the captured SynchronizationContext, and the latter just served as a simple factory to do that capture and create the AsyncOperation instance. Then EAP implementations would use those, e.g. Ping.SendAsync called [AsyncOperationManager.CreateOperation](https://github.com/dotnet/runtime/blob/5f94bffeff62f4b767a311a4505d6d40d86279d9/src/libraries/System.ComponentModel.EventBasedAsync/src/System/ComponentModel/AsyncOperationManager.cs#L10-L36) to capture the SynchronizationContext, and then when the operation completed, the AsyncOperations [PostOperationCompleted](https://github.com/dotnet/runtime/blob/5f94bffeff62f4b767a311a4505d6d40d86279d9/src/libraries/System.ComponentModel.EventBasedAsync/src/System/ComponentModel/AsyncOperation.cs#L51-L77) method would be invoked to call the stored SynchronizationContexts Post method.
SynchronizationContext provides a few more trinkets worthy of mention as theyll show up again in a bit. In particular, it exposes OperationStarted and OperationCompleted methods. The base implementation of these virtuals are empty, doing nothing, but a derived implementation might override these to know about in-flight operations. That means EAP implementations would also invoke these OperationStarted/OperationCompleted at the beginning and end of each operation, in order to inform any present SynchronizationContext and allow it to track the work. This is particularly relevant to the EAP pattern because the methods that initiate the async operations are void returning: you get nothing back that allows you to track the work individually. Well get back to that.
So, we needed something better than the APM pattern, and the EAP that came next introduced some new things but didnt really address the core problems we faced. We still needed something better.
## Enter Tasks
.NET Framework 4.0 introduced the System.Threading.Tasks.Task type. At its heart, a Task is just a data structure that represents the eventual completion of some asynchronous operation (other frameworks call a similar type a “promise” or a “future”). A Task is created to represent some operation, and then when the operation it logically represents completes, the results are stored into that Task. Simple enough. But _the_ key feature that Task provides that makes it leaps and bounds more useful than IAsyncResult is that it builds into itself the notion of a continuation. That one feature means you can walk up to any Task and ask to be notified asynchronously when it completes, with the task itself handling the synchronization to ensure the continuation is invoked regardless of whether the task has already completed, hasnt yet completed, or is completing concurrently with the notification request. Why is that so impactful? Well, if you remember back to our discussion of the old APM pattern, there were two primary problems.
1. You had to implement a custom IAsyncResult implementation for every operation: there was no built-in IAsyncResult implementation anyone could just use for their needs.
2. You had to know prior to the Begin method being called what you wanted to do when it was complete. This makes it a significant challenge to implement combinators and other generalized routines for consuming and composing arbitrary async implementations.
In contrast, with Task, that shared representation lets you walk up to an async operation _after_ youve already initiated the operation and provide a continuation _after_ youve already initiated the operation… you dont need to provide that continuation _to_ the method that initiates the operation. Everyone who has asynchronous operations can produce a Task, and everyone who consumes asynchronous operations can consume a Task, and nothing custom needs to be done to connect the two: Task becomes the lingua franca for enabling producers and consumers of asynchronous operations to talk. And that has changed the face of .NET. More on that in a bit…
For now, lets better understand what this actually means. Rather than dive into the intricate code for Task, well do the pedagogical thing and just implement a simple version. This isnt meant to be a great implementation, rather only complete enough functionally to help understand the meat of what is a Task, which, at the end of the day, is really just a data structure that handles coordinating the setting and reception of a completion signal. Well start with just a few fields:
class MyTask{ private bool _completed; private Exception? _error; private Action<MyTask>? _continuation; private ExecutionContext? _ec; ...}
We need a field to know whether the task has completed (_completed), and we need a field to store any error that caused the task to fail (_error); if we were also implementing a generic MyTask<TResult>, thered also be a private TResult _result field for storing the successful result of the operation. Thus far, this looks a lot like our custom IAsyncResult implementation earlier (not a coincidence, of course). But now the pièce de résistance, the _continuation field. In this simple implementation, were supporting just a single continuation, but thats enough for explanatory purposes (the real Task employs an [object](https://github.com/dotnet/runtime/blob/81977309048600e67fdb44a7d4c99aaad89846d7/src/libraries/System.Private.CoreLib/src/System/Threading/Tasks/Task.cs#L176-L178) field that can either be an individual continuation object or a List<> of continuation objects). This is a delegate that will be invoked when the task completes.
Now, a bit of surface area. As noted, one of the fundamental advances in Task over previous models was the ability to supply the continuation work (the callback) _after_ the operation was initiated. We need a method to let us do that, so lets add ContinueWith:
public void ContinueWith(Action<MyTask> action){ lock (this) { if (_completed) { ThreadPool.QueueUserWorkItem(_ => action(this)); } else if (_continuation is not null) { throw new InvalidOperationException("Unlike Task, this implementation only supports a single continuation."); } else { _continuation = action; _ec = ExecutionContext.Capture(); } }}
If the task has already been marked completed by the time ContinueWith is called, ContinueWith just queues the execution of the delegate. Otherwise, the method stores the delegate, such that the continuation may be queued when the task completes (it also stores something called an ExecutionContext, and then uses that when the delegate is later invoked, but dont worry about that part for now… well get to it). Simple enough.
Then we need to be able to mark the MyTask as completed, meaning whatever asynchronous operation it represents has finished. For that, well expose two methods, one to mark it completed successfully (“SetResult”), and one to mark it completed with an error (“SetException”):
public void SetResult() => Complete(null);public void SetException(Exception error) => Complete(error);private void Complete(Exception? error){ lock (this) { if (_completed) { throw new InvalidOperationException("Already completed"); } _error = error; _completed = true; if (_continuation is not null) { ThreadPool.QueueUserWorkItem(_ => { if (_ec is not null) { ExecutionContext.Run(_ec, _ => _continuation(this), null); } else { _continuation(this); } }); } }}
We store any error, we mark the task as having been completed, and then if a continuation had previously been registered, we queue it to be invoked.
Finally, we need a way to propagate any exception that may have occurred in the task (and, if this were a generic MyTask<T>, to return its _result); to facilitate certain scenarios, we also allow this method to block waiting for the task to complete, which we can implement in terms of ContinueWith (the continuation just signals a ManualResetEventSlim that the caller then blocks on waiting for completion).
public void Wait(){ ManualResetEventSlim? mres = null; lock (this) { if (!_completed) { mres = new ManualResetEventSlim(); ContinueWith(_ => mres.Set()); } } mres?.Wait(); if (_error is not null) { ExceptionDispatchInfo.Throw(_error); }}
And thats basically it. Now to be sure, the real Task is way more complicated, with a much more efficient implementation, with support for any number of continuations, with a multitude of knobs about how it should behave (e.g. should continuations be queued as is being done here or should they be invoked synchronously as part of the tasks completion), with the ability to store multiple exceptions rather than just one, with special knowledge of cancellation, with tons of helper methods for doing common operations (e.g. Task.Run which creates a Task to represent a delegate queued to be invoked on the thread pool), and so on. But theres no magic to any of that; at its core, its just what we saw here.
You might also notice that my simple MyTask has public SetResult/SetException methods directly on it, whereas Task doesnt. Actually, Task _does_ have such methods, [theyre just internal](https://github.com/dotnet/runtime/blob/81977309048600e67fdb44a7d4c99aaad89846d7/src/libraries/System.Private.CoreLib/src/System/Threading/Tasks/Task.cs#L3271), with a System.Threading.Tasks.TaskCompletionSource type serving as a separate “producer” for the task and its completion; that was done not out of technical necessity but as a way to keep the completion methods off of the thing meant only for consumption. You can then hand out a Task without having to worry about it being completed out from under you; the completion signal is an implementation detail of whatever created the task and also reserves the right to complete it by keeping the TaskCompletionSource to itself. (CancellationToken and CancellationTokenSource follow a similar pattern: CancellationToken is just a struct wrapper for a CancellationTokenSource, serving up only the public surface area related to consuming a cancellation signal but without the ability to produce one, which is a capability restricted to whomever has access to the CancellationTokenSource.)
Of course, we can implement combinators and helpers for this MyTask similar to what Task provides. Want a simple MyTask.WhenAll? Here you go:
public static MyTask WhenAll(MyTask t1, MyTask t2){ var t = new MyTask(); int remaining = 2; Exception? e = null; Action<MyTask> continuation = completed => { e ??= completed._error; // just store a single exception for simplicity if (Interlocked.Decrement(ref remaining) == 0) { if (e is not null) t.SetException(e); else t.SetResult(); } }; t1.ContinueWith(continuation); t2.ContinueWith(continuation); return t;}
Want a MyTask.Run? You got it:
public static MyTask Run(Action action){ var t = new MyTask(); ThreadPool.QueueUserWorkItem(_ => { try { action(); t.SetResult(); } catch (Exception e) { t.SetException(e); } }); return t;}
How about a MyTask.Delay? Sure:
public static MyTask Delay(TimeSpan delay){ var t = new MyTask(); var timer = new Timer(_ => t.SetResult()); timer.Change(delay, Timeout.InfiniteTimeSpan); return t;}
You get the idea.
With Task in place, all previous async patterns in .NET became a thing of the past. Anywhere an asynchronous implementation previously was implemented with the APM pattern or the EAP pattern, new Task-returning methods were exposed.
### And ValueTasks
Task continues to be the workhorse for asynchrony in .NET to this day, with new methods exposed every release and routinely throughout the ecosystem that return Task and Task<TResult>. However, Task is a class, which means creating one does come with an allocation. For the most part, one extra allocation for a long-lived asynchronous operation is a pittance and wont meaningfully impact performance for all but the most performance-sensitive operations. However, as was previously noted, synchronous completion of asynchronous operations is fairly common. Stream.ReadAsync was introduced to return a Task<int>, but if youre reading from, say, a BufferedStream, theres a really good chance many of your reads are going to complete synchronously due to simply needing to pull data from an in-memory buffer rather than performing syscalls and real I/O. Having to allocate an additional object just to return such data is unfortunate (note it was the case with APM as well). For non-generic Task-returning methods, the method can just return a singleton already-completed task, and in fact one such singleton is provided by Task in the form of Task.CompletedTask. But for Task<TResult>, its impossible to cache a Task for every possible TResult. What can we do to make such synchronous completion faster?
It is possible to cache _some_ Task<TResult>s. For example, Task<bool> is very common, and theres only two meaningful things to cache there: a Task<bool> when the Result is true and one when the Result is false. Or while we wouldnt want to try caching four billion Task<int>s to accommmodate every possible Int32 result, small Int32 values are very common, so we could cache a few for, say, -1 through 8. Or for arbitrary types, default is a reasonably common value, so we could cache a Task<TResult> where Result is default(TResult) for every relevant type. And in fact, [Task.FromResult](https://github.com/dotnet/runtime/blob/81977309048600e67fdb44a7d4c99aaad89846d7/src/libraries/System.Private.CoreLib/src/System/Threading/Tasks/Task.cs#L5222-L5273) does that today (as of recent versions of .NET), using a small cache of such reusable Task<TResult> singletons and returning one of them if appropriate or otherwise allocating a new Task<TResult> for the exact provided result value. Other schemes can be created to handle other reasonably common cases. For example, when working with Stream.ReadAsync, its reasonably common to call it multiple times on the same stream, all with the same count for the number of bytes allowed to be read. And its reasonably common for the implementation to be able to fully satisfy that count request. Which means its reasonably common for Stream.ReadAsync to repeatedly return the same int result value. To avoid multiple allocations in such scenarios, multiple Stream types (like MemoryStream) will cache the last Task<int> they successfully returned, and if the next read ends up also completing synchronously and successfully with the same result, it can just return the same Task<int> again rather than creating a new one. But what about other cases? How can this allocation for synchronous completions be avoided more generally in situations where the performance overhead really matters?
Thats where ValueTask<TResult> comes into the picture ([a much more detailed examination of](https://devblogs.microsoft.com/dotnet/understanding-the-whys-whats-and-whens-of-valuetask/) ValueTask<TResult> is also available). ValueTask<TResult> started life as a discriminated union between a TResult and a Task<TResult>. At the end of the day, ignoring all the bells and whistles, [thats all it is](https://github.com/dotnet/corefx/blob/d6173e069a9bcedfdfd7f4f41e67d23f67157b61/src/System.Threading.Tasks.Extensions/src/System/Threading/Tasks/ValueTask.cs#L53-L58) (or, rather, was), either an immediate result or a promise for a result at some point in the future:
public readonly struct ValueTask<TResult>{ private readonly Task<TResult>? _task; private readonly TResult _result; ...}
A method could then return such a ValueTask<TResult> instead of a Task<TResult>, and at the expense of a larger return type and a little more indirection, avoid the Task<TResult> allocation if the TResult was known by the time it needed to be returned.
There are, however, super duper extreme high-performance scenarios where you want to be able to avoid the Task<TResult> allocation even in the asynchronous-completion case. For example, Socket lives at the bottom of the networking stack, and SendAsync and ReceiveAsync on sockets are on the super hot path for many a service, with both synchronous and asynchronous completions being very common (most sends complete synchronously, and many receives complete synchronously due to data having already been buffered in the kernel). Wouldnt it be nice if, on a given Socket, we could make such sending and receiving allocation-free, regardless of whether the operations complete synchronously or asynchronously?
Thats where System.Threading.Tasks.Sources.IValueTaskSource<TResult> enters the picture:
public interface IValueTaskSource<out TResult>{ ValueTaskSourceStatus GetStatus(short token); void OnCompleted(Action<object?> continuation, object? state, short token, ValueTaskSourceOnCompletedFlags flags); TResult GetResult(short token);}
The IValueTaskSource<TResult> interface allows an implementation to provide its own backing object for a ValueTask<TResult>, enabling the object to implement methods like GetResult to retrieve the result of the operation and OnCompleted to hook up a continuation to the operation. With that, ValueTask<TResult> evolved [a small change to its definition](https://github.com/dotnet/runtime/blob/81977309048600e67fdb44a7d4c99aaad89846d7/src/libraries/System.Private.CoreLib/src/System/Threading/Tasks/ValueTask.cs#L465-L468), with its Task<TResult>? _task field replaced by an object? _obj field:
public readonly struct ValueTask<TResult>{ private readonly object? _obj; private readonly TResult _result; ...}
Whereas the _task field was either a Task<TResult> or null, the _obj field now can also be an IValueTaskSource<TResult>. Once a Task<TResult> is marked as completed, thats it, it will remain completed and never transition back to an incomplete state. In contrast, an object implementing IValueTaskSource<TResult> has full control over the implementation, and is free to transition bidirectionally between complete and incomplete states, as ValueTask<TResult>s contract is that a given instance may be consumed only once, thus by construction it shouldnt observe a post-consumption change in the underlying instance (this is why analysis rules like [CA2012](https://learn.microsoft.com/dotnet/fundamentals/code-analysis/quality-rules/ca2012) exist). This then enables types like Socket to pool IValueTaskSource<TResult> instances to use for repeated calls. Socket caches up to two such instances, one for reads and one for writes, since the 99.999% case is to have at most one receive and one send in-flight at the same time.
I mentioned ValueTask<TResult> but not ValueTask. When dealing only with avoiding allocation for synchronous completion, theres little performance benefit to a non-generic ValueTask (representing result-less, void operations), since the same condition can be represented with Task.CompletedTask. But once we care about the ability to use a poolable underlying object for avoiding allocation in asynchronous completion case, that then also matters for the non-generic. Thus, when IValueTaskSource<TResult> was introduced, so too were IValueTaskSource and ValueTask.
So, we have Task, Task<TResult>, ValueTask, and ValueTask<TResult>. Were able to interact with them in various ways, representing arbitrary asynchronous operations and hooking up continuations to handle the completion of those asynchronous operations. And yes, we can do so _before_ or _after_ the operation completes.
_But_… those continuations are still callbacks!
Were still forced into a continuation-passing style for encoding our asynchronous control flow!!
Its still really hard to get right!!!
How can we fix that????
## C# Iterators to the Rescue
The glimmer of hope for that solution actually came about a few years before Task hit the scene, with C# 2.0, when it added support for iterators.
“Iterators?” you ask? “You mean for IEnumerable<T>?” Thats the one. Iterators let you write a single method that is then used by the compiler to implement an IEnumerable<T> and/or an IEnumerator<T>. For example, if I wanted to create an enumerable that yielded the Fibonacci sequence, I might write something like this:
public static IEnumerable<int> Fib(){ int prev = 0, next = 1; yield return prev; yield return next; while (true) { int sum = prev + next; yield return sum; prev = next; next = sum; }}
I can then enumerate this with a foreach:
foreach (int i in Fib()){ if (i > 100) break; Console.Write($"{i} ");}
I can compose it with other IEnumerable<T>s via combinators like those on System.Linq.Enumerable:
foreach (int i in Fib().Take(12)){ Console.Write($"{i} ");}
Or I can just manually enumerate it directly via an IEnumerator<T>:
using IEnumerator<int> e = Fib().GetEnumerator();while (e.MoveNext()){ int i = e.Current; if (i > 100) break; Console.Write($"{i} ");}
All of the above result in this output:
0 1 1 2 3 5 8 13 21 34 55 89
The really interesting thing about this is that in order to achieve the above, we need to be able to enter and exit that Fib method multiple times. We call MoveNext, it enters the method, the method then executes until it encounters a yield return, at which point the call to MoveNext needs to return true and a subsequent access to Current needs to return the yielded value. Then we call MoveNext again, and we need to be able to pick up in Fib just after where we last left off, and with all of the state from the previous invocation intact. Iterators are effectively coroutines provided by the C# language / compiler, with the compiler expanding my Fib iterator into a full-blown state machine:
public static IEnumerable<int> Fib() => new <Fib>d__0(-2);[CompilerGenerated]private sealed class <Fib>d__0 : IEnumerable<int>, IEnumerable, IEnumerator<int>, IEnumerator, IDisposable{ private int <>1__state; private int <>2__current; private int <>l__initialThreadId; private int <prev>5__2; private int <next>5__3; private int <sum>5__4; int IEnumerator<int>.Current => <>2__current; object IEnumerator.Current => <>2__current; public <Fib>d__0(int <>1__state) { this.<>1__state = <>1__state; <>l__initialThreadId = Environment.CurrentManagedThreadId; } private bool MoveNext() { switch (<>1__state) { default: return false; case 0: <>1__state = -1; <prev>5__2 = 0; <next>5__3 = 1; <>2__current = <prev>5__2; <>1__state = 1; return true; case 1: <>1__state = -1; <>2__current = <next>5__3; <>1__state = 2; return true; case 2: <>1__state = -1; break; case 3: <>1__state = -1; <prev>5__2 = <next>5__3; <next>5__3 = <sum>5__4; break; } <sum>5__4 = <prev>5__2 + <next>5__3; <>2__current = <sum>5__4; <>1__state = 3; return true; } IEnumerator<int> IEnumerable<int>.GetEnumerator() { if (<>1__state == -2 && <>l__initialThreadId == Environment.CurrentManagedThreadId) { <>1__state = 0; return this; } return new <Fib>d__0(0); } IEnumerator IEnumerable.GetEnumerator() => ((IEnumerable<int>)this).GetEnumerator(); void IEnumerator.Reset() => throw new NotSupportedException(); void IDisposable.Dispose() { }}
All of the logic for Fib is now inside of the MoveNext method, but as part of a jump table that lets the implementation branch to where it last left off, which is tracked in a generated state field on the enumerator type. And the variables I wrote as locals, like prev, next, and sum, have been “lifted” to be fields on the enumerator, so that they may persist across invocations of MoveNext.
(Note that the previous code snippet showing how the C# compiler emits the implementation wont compile as-is. The C# compiler synthesizes “unspeakable” names, meaning it names types and members it creates in a way thats valid IL but invalid C#, so as not to risk conflicting with any user-named types and members. Ive kept everything named as the compiler does, but if you want to experiment with compiling it, you can rename things to use valid C# names instead.)
In my previous example, the last form of enumeration I showed involved manually using the IEnumerator<T>. At that level, were manually invoking MoveNext(), deciding when it was an appropriate time to re-enter the coroutine. But… what if instead of invoking it like that, I could instead have the next invocation of MoveNext actually be part of the continuation work performed when an asynchronous operation completes? What if I could yield return something that represents an asynchronous operation and have the consuming code hook up a continuation to that yielded object where that continuation then does the MoveNext? With such an approach, I could write a helper method like this:
static Task IterateAsync(IEnumerable<Task> tasks){ var tcs = new TaskCompletionSource(); IEnumerator<Task> e = tasks.GetEnumerator(); void Process() { try { if (e.MoveNext()) { e.Current.ContinueWith(t => Process()); return; } } catch (Exception e) { tcs.SetException(e); return; } tcs.SetResult(); }; Process(); return tcs.Task;}
Now this is getting interesting. Were given an enumerable of tasks that we can iterate through. Each time we MoveNext to the next Task and get one, we then hook up a continuation to that Task; when that Task completes, itll just turn around and call right back to the same logic that does a MoveNext, gets the next Task, and so on. This is building on the idea of Task as a single representation for any asynchronous operation, so the enumerable were fed can be a sequence of any asynchronous operations. Where might such a sequence come from? From an iterator, of course. Remember our earlier CopyStreamToStream example and how gloriously horrible the APM-based implementation was? Consider this instead:
static Task CopyStreamToStreamAsync(Stream source, Stream destination){ return IterateAsync(Impl(source, destination)); static IEnumerable<Task> Impl(Stream source, Stream destination) { var buffer = new byte[0x1000]; while (true) { Task<int> read = source.ReadAsync(buffer, 0, buffer.Length); yield return read; int numRead = read.Result; if (numRead <= 0) { break; } Task write = destination.WriteAsync(buffer, 0, numRead); yield return write; write.Wait(); } }}
Wow, this is almost legible. Were calling that IterateAsync helper, and the enumerable were feeding it is one produced by an iterator thats handling all the control flow for the copy. It calls Stream.ReadAsync and then yield returns that Task; that yielded task is what will be handed off to IterateAsync after it calls MoveNext, and IterateAsync will hook a continuation up to that Task, which when it completes will then just call back into MoveNext and end up back in this iterator just after the yield. At that point, the Impl logic gets the result of the method, calls WriteAsync, and again yields the Task it produced. And so on.
And that, my friends, is the beginning of async/await in C# and .NET. Something around 95% of the logic in support of iterators and async/await in the C# compiler is shared. Different syntax, different types involved, but fundamentally the same transform. Squint at the yield returns, and you can almost see awaits in their stead.
In fact, some enterprising developers [used iterators in this fashion for asynchronous programming](https://learn.microsoft.com/archive/msdn-magazine/2008/june/concurrent-affairs-simplified-apm-with-the-asyncenumerator) before async/await hit the scene. And a similar transformation was prototyped in the experimental [Axum](https://en.wikipedia.org/wiki/Axum_(programming_language)) programming language, serving as a key inspiration for C#s async support. Axum provided an async keyword that could be put onto a method, just like async can now in C#. Task wasnt yet ubiquitous, so inside of async methods, the Axum compiler heuristically matched synchronous method calls to their APM counterparts, e.g. if it saw you calling stream.Read, it would find and utilize the corresponding stream.BeginRead and stream.EndRead methods, synthesizing the appropriate delegate to pass to the Begin method, while also generating a complete APM implementation for the async method being defined such that it was compositional. It even integrated with SynchronizationContext! While Axum was eventually shelved, it served as an awesome and motivating prototype for what eventually became async/await in C#.
## async/await under the covers
Now that we know how we got here, lets dive in to how it actually works. For reference, heres our example synchronous method again:
public void CopyStreamToStream(Stream source, Stream destination){ var buffer = new byte[0x1000]; int numRead; while ((numRead = source.Read(buffer, 0, buffer.Length)) != 0) { destination.Write(buffer, 0, numRead); }}
and again heres what the corresponding method looks like with async/await:
public async Task CopyStreamToStreamAsync(Stream source, Stream destination){ var buffer = new byte[0x1000]; int numRead; while ((numRead = await source.ReadAsync(buffer, 0, buffer.Length)) != 0) { await destination.WriteAsync(buffer, 0, numRead); }}
A breadth of fresh air in comparison to everything weve seen thus far. The signature changed from void to async Task, we call ReadAsync and WriteAsync instead of Read and Write, respectively, and both of those operations are prefixed with await. Thats it. The compiler and the core libraries take over the rest, fundamentally changing how the code is actually executed. Lets dive into how.
### Compiler Transform
As weve already seen, as with iterators, the compiler rewrites the async method into one based on a state machine. We still have a method with the same signature the developer wrote (public Task CopyStreamToStreamAsync(Stream source, Stream destination)), but the body of that method is completely different:
[AsyncStateMachine(typeof(<CopyStreamToStreamAsync>d__0))]public Task CopyStreamToStreamAsync(Stream source, Stream destination){ <CopyStreamToStreamAsync>d__0 stateMachine = default; stateMachine.<>t__builder = AsyncTaskMethodBuilder.Create(); stateMachine.source = source; stateMachine.destination = destination; stateMachine.<>1__state = -1; stateMachine.<>t__builder.Start(ref stateMachine); return stateMachine.<>t__builder.Task;}private struct <CopyStreamToStreamAsync>d__0 : IAsyncStateMachine{ public int <>1__state; public AsyncTaskMethodBuilder <>t__builder; public Stream source; public Stream destination; private byte[] <buffer>5__2; private TaskAwaiter <>u__1; private TaskAwaiter<int> <>u__2; ...}
Note that the only signature difference from what the dev wrote is the lack of the async keyword itself. async isnt actually a part of the method signature; like unsafe, when you put it in the method signature, youre expressing an implementation detail of the method rather than something thats actually exposed as part of the contract. Using async/await to implement a Task-returning method is an implementation detail.
The compiler has generated a struct named <CopyStreamToStreamAsync>d__0, and its zero-initialized an instance of that struct on the stack. Importantly, if the async method completes synchronously, this state machine will never have left the stack. That means theres no allocation associated with the state machine _unless_ the method needs to complete asynchronously, meaning it awaits something thats not yet completed by that point. More on that in a bit.
This struct _is_ the state machine for the method, containing not only all of the transformed logic from what the developer wrote, but also fields for tracking the current position in that method as well as all of the “local” state the compiler lifted out of the method that needs to survive between MoveNext invocations. Its the logical equivalent of the IEnumerable<T>/IEnumerator<T> implementation we saw in the iterator. (Note that the code Im showing is from a release build; in debug builds the C# compiler will actually generate these state machine types as classes, as doing so can aid in certain debugging exercises).
After initializing the state machine, we see a call to [AsyncTaskMethodBuilder.Create()](https://github.com/dotnet/runtime/blob/6319039691477bf9296a0d62fd4a2491868966d8/src/libraries/System.Private.CoreLib/src/System/Runtime/CompilerServices/AsyncTaskMethodBuilder.cs#L25). While were currently focused on Tasks, the C# language and compiler allow for arbitrary types ([“task-like” types](https://learn.microsoft.com/dotnet/csharp/language-reference/proposals/csharp-7.0/task-types#builder-type)) to be returned from async methods, e.g. I can write a method public async MyTask CopyStreamToStreamAsync, and it would compile just fine as long as we augment the MyTask we defined earlier in an appropriate way. That appropriateness includes declaring an associated “builder” type and associating it with the type via the AsyncMethodBuilder attribute:
[AsyncMethodBuilder(typeof(MyTaskMethodBuilder))]public class MyTask{ ...}public struct MyTaskMethodBuilder{ public static MyTaskMethodBuilder Create() { ... } public void Start<TStateMachine>(ref TStateMachine stateMachine) where TStateMachine : IAsyncStateMachine { ... } public void SetStateMachine(IAsyncStateMachine stateMachine) { ... } public void SetResult() { ... } public void SetException(Exception exception) { ... } public void AwaitOnCompleted<TAwaiter, TStateMachine>( ref TAwaiter awaiter, ref TStateMachine stateMachine) where TAwaiter : INotifyCompletion where TStateMachine : IAsyncStateMachine { ... } public void AwaitUnsafeOnCompleted<TAwaiter, TStateMachine>( ref TAwaiter awaiter, ref TStateMachine stateMachine) where TAwaiter : ICriticalNotifyCompletion where TStateMachine : IAsyncStateMachine { ... } public MyTask Task { get { ... } }}
In this context, such a “builder” is something that knows how to create an instance of that type (the Task property), complete it either successfully and with a result if appropriate (SetResult) or with an exception (SetException), and handle hooking up continuations to awaited things that havent yet completed (AwaitOnCompleted/AwaitUnsafeOnCompleted). In the case of System.Threading.Tasks.Task, it is by default associated with the AsyncTaskMethodBuilder. Normally that association is provided via an [AsyncMethodBuilder(...)] attribute applied to the type, but Task is known specially to C# and so isnt actually adorned with that attribute. As such, the compiler has reached for the builder to use for this async method, and is constructing an instance of it using the Create method thats part of the pattern. Note that as with the state machine, AsyncTaskMethodBuilder is also a struct, so theres no allocation here, either.
The state machine is then populated with the arguments to this entry point method. Those parameters need to be available to the body of the method thats been moved into MoveNext, and as such these arguments need to be stored in the state machine so that they can be referenced by the code on the subsequent call to MoveNext. The state machine is also initialized to be in the initial -1 state. If MoveNext is called and the state is -1, well end up starting logically at the beginning of the method.
Now the most unassuming but most consequential line: a call to the builders [Start](https://github.com/dotnet/runtime/blob/6319039691477bf9296a0d62fd4a2491868966d8/src/libraries/System.Private.CoreLib/src/System/Runtime/CompilerServices/AsyncTaskMethodBuilder.cs#L32-L33) method. This is another part of the pattern that must be exposed on a type used in the return position of an async method, and its used to perform the initial MoveNext on the state machine. The builders Start method is effectively just this:
public void Start<TStateMachine>(ref TStateMachine stateMachine) where TStateMachine : IAsyncStateMachine{ stateMachine.MoveNext();}
such that calling stateMachine.<>t__builder.Start(ref stateMachine); is really just calling stateMachine.MoveNext(). In which case, why doesnt the compiler just emit that directly? Why have Start at all? The answer is that theres a tad bit more to Start than I let on. But for that, we need to take a brief detour into understanding ExecutionContext.
_ExecutionContext_
Were all familiar with passing around state from method to method. You call a method, and if that method specifies parameters, you call the method with arguments in order to feed that data into the callee. This is explicitly passing around data. But there are other more implicit means. For example, rather than passing data as arguments, a method could be parameterless but could dictate that some specific static fields may be populated prior to making the method call, and the method will pull state from there. Nothing about the methods signature indicates it takes arguments, because it doesnt: theres just an implicit contract between the caller and callee that the caller might populate some memory locations and the callee might read those memory locations. The callee and the caller may not even realize its happening if theyre intermediaries, e.g. method A might populate the statics and then call B which calls C which calls D which eventually calls E that reads the values of those statics. This is often referred to as “ambient” data: its not passed to you via parameters but rather is just sort of hanging out there and available for you to consume if desired.
We can take this a step further, and use thread-local state. Thread-local state, which in .NET is achieved via static fields attributed as [ThreadStatic] or via the ThreadLocal<T> type, can be used in the same way, but with the data limited to just the current thread of execution, with every thread able to have its own isolated copy of those fields. With that, you could populate the thread static, make the method call, and then upon the methods completion revert the changes to the thread static, enabling a fully isolated form of such implicitly passed data.
But, what about asynchrony? If we make an asynchronous method call and logic inside that asynchronous method wants to access that ambient data, how would it do so? If the data were stored in regular statics, the asynchronous method would be able to access it, but you could only ever have one such method in flight at a time, as multiple callers could end up overwriting each others state when they write to those shared static fields. If the data were stored in thread statics, the asynchronous method would be able to access it, but only up until the point where it stopped running synchronously on the calling thread; if it hooked up a continuation to some operation it initiated and that continuation ended up running on some other thread, it would no longer have access to the thread static information. Even if it did happen to run on the same thread, either by chance or because the scheduler forced it to, by the time it did its likely the data would have been removed and/or overwritten by some other operation initiated by that thread. For asynchrony, what we need is a mechanism that would allow arbitrary ambient data to flow across these asynchronous points, such that throughout an async methods logic, wherever and whenever that logic might run, it would have access to that same data.
Enter ExecutionContext. The ExecutionContext type is the vehicle by which ambient data flows from async operation to async operation. It lives in a [ThreadStatic], but then when some asynchronous operation is initiated, its “captured” (a fancy way of saying “read a copy from that thread static”), stored, and then when the continuation of that asynchronous operation is run, the ExecutionContext is first restored to live in the [ThreadStatic] on the thread which is about to run the operation. ExecutionContext is the mechanism by which AsyncLocal<T> is implemented (in fact, in .NET Core, ExecutionContext is entirely about AsyncLocal<T>, nothing more), such that if you store a value into an AsyncLocal<T>, and then for example queue a work item to run on the ThreadPool, that value will be visible in that AsyncLocal<T> inside of that work item running on the pool:
var number = new AsyncLocal<int>();number.Value = 42;ThreadPool.QueueUserWorkItem(_ => Console.WriteLine(number.Value));number.Value = 0;Console.ReadLine();
That will print 42 every time this is run. It doesnt matter that the moment after we queue the delegate we reset the value of the AsyncLocal<int> back to 0, because the ExecutionContext was captured as part of the QueueUserWorkItem call, and that capture included the state of the AsyncLocal<int> at that exact moment. We can see this in more detail by implementing our own simple thread pool:
using System.Collections.Concurrent;var number = new AsyncLocal<int>();number.Value = 42;MyThreadPool.QueueUserWorkItem(() => Console.WriteLine(number.Value));number.Value = 0;Console.ReadLine();class MyThreadPool{ private static readonly BlockingCollection<(Action, ExecutionContext?)> s_workItems = new(); public static void QueueUserWorkItem(Action workItem) { s_workItems.Add((workItem, ExecutionContext.Capture())); } static MyThreadPool() { for (int i = 0; i < Environment.ProcessorCount; i++) { new Thread(() => { while (true) { (Action action, ExecutionContext? ec) = s_workItems.Take(); if (ec is null) { action(); } else { ExecutionContext.Run(ec, s => ((Action)s!)(), action); } } }) { IsBackground = true }.UnsafeStart(); } }}
Here MyThreadPool has a BlockingCollection<(Action, ExecutionContext?)> that represents its work item queue, with each work item being the delegate for the work to be invoked as well as the ExecutionContext associated with that work. The static constructor for the pool spins up a bunch of threads, each of which just sits in an infinite loop taking the next work item and running it. If no ExecutionContext was captured for a given delegate, the delegate is just invoked directly. But if an ExecutionContext was captured, rather than invoking the delegate directly, we call the ExecutionContext.Run method, which will restore the supplied ExecutionContext as the current context prior to running the delegate, and will then reset the context afterwards. This example includes the exact same code with an AsyncLocal<int> previously shown, except this time using MyThreadPool instead of ThreadPool, yet it will still output 42 each time, because the pool is properly flowing ExecutionContext.
As an aside, youll note I called UnsafeStart in MyThreadPools static constructor. Starting a new thread is exactly the kind of asynchronous point that should flow ExecutionContext, and indeed, Threads Start method uses ExecutionContext.Capture to capture the current context, store it on the Thread, and then use that captured context when eventually invoking the Threads ThreadStart delegate. I didnt want to do that in this example, though, as I didnt want the Threads to capture whatever ExecutionContext happened to be present when the static constructor ran (doing so could make a demo about ExecutionContext more convoluted), so I used the UnsafeStart method instead. Threading-related methods that begin with Unsafe behave exactly the same as the corresponding method that lacks the Unsafe prefix except that they _dont_ capture ExecutionContext, e.g. Thread.Start and Thread.UnsafeStart do identical work, but whereas Start captures ExecutionContext, UnsafeStart does not.
_Back To Start_
We took a detour into discussing ExecutionContext when I was writing about the implementation of AsyncTaskMethodBuilder.Start, which I said was effectively:
public void Start<TStateMachine>(ref TStateMachine stateMachine) where TStateMachine : IAsyncStateMachine{ stateMachine.MoveNext();}
and then suggested I simplified a bit. That simplification was ignoring the fact that the method actually needs to factor ExecutionContext into things, and is thus more like this:
public void Start<TStateMachine>(ref TStateMachine stateMachine) where TStateMachine : IAsyncStateMachine{ ExecutionContext previous = Thread.CurrentThread._executionContext; // [ThreadStatic] field try { stateMachine.MoveNext(); } finally { ExecutionContext.Restore(previous); // internal helper }}
Rather than just calling stateMachine.MoveNext() as Id previously suggested we did, we do a dance here of getting the current ExecutionContext, then invoking MoveNext, and then upon its completion resetting the current context back to what it was prior to the MoveNext invocation.
The reason for this is to prevent ambient data leakage from an async method out to its caller. An example method demonstrates why that matters:
async Task ElevateAsAdminAndRunAsync(){ using (WindowsIdentity identity = LoginAdmin()) { using (WindowsImpersonationContext impersonatedUser = identity.Impersonate()) { await DoSensitiveWorkAsync(); } }}
“Impersonation” is the act of changing ambient information about the current user to instead be that of someone else; this lets code act on behalf of someone else, using their privileges and access. In .NET, such impersonation flows across asynchronous operations, which means its part of ExecutionContext. Now imagine if Start didnt restore the previous context, and consider this code:
Task t = ElevateAsAdminAndRunAsync();PrintUser();await t;
This code could find that the ExecutionContext modified inside of ElevateAsAdminAndRunAsync remains after ElevateAsAdminAndRunAsync returns to its synchronous caller (which happens the first time the method awaits something thats not yet complete). Thats because after calling Impersonate, we call DoSensitiveWorkAsync and await the task it returns. Assuming that task isnt complete, it will cause the invocation of ElevateAsAdminAndRunAsync to yield and return to the caller, with the impersonation still in effect on the current thread. That is not something we want. As such, Start erects this guard that ensures any modifications to ExecutionContext dont flow _out_ of the synchronous method call and only flow along with any subsequent work performed by the method.
_MoveNext_
So, the entry point method was invoked, the state machine struct was initialized, Start was called, and that invoked MoveNext. What is MoveNext? Its the method that contains all of the original logic from the devs method, but with a whole bunch of changes. Lets start just by looking at the scaffolding of the method. Heres a decompiled version of what the compiler emit for our method, but with everything inside of the generated try block removed:
private void MoveNext(){ try { ... // all of the code from the CopyStreamToStreamAsync method body, but not exactly as it was written } catch (Exception exception) { <>1__state = -2; <buffer>5__2 = null; <>t__builder.SetException(exception); return; } <>1__state = -2; <buffer>5__2 = null; <>t__builder.SetResult();}
Whatever other work is performed by MoveNext, it has the responsibility of completing the Task returned from the async Task method when all of the work is done. If the body of the try block throws an exception that goes unhandled, then the task will be faulted with that exception. And if the async method successfully reaches its end (equivalent to a synchronous method returning), it will complete the returned task successfully. In either of those cases, its setting the state of the state machine to indicate completion. (I sometimes hear developers theorize that, when it comes to exceptions, theres a difference between those thrown before the first await and after… based on the above, it should be clear that is _not_ the case. Any exception that goes unhandled inside of an async method, no matter where it is in the method and no matter whether the method has yielded, will end up in the above catch block, with the caught exception then stored into the Task thats returned from the async method.)
Also note that this completion is going through the builder, using its SetException and SetResult methods that are part of the pattern for a builder expected by the compiler. If the async method has previously suspended, the builder will have already had to manufacture a Task as part of that suspension handling (well see how and where soon), in which case calling SetException/SetResult will complete that Task. If, however, the async method hasnt previously suspended, then we havent yet created a Task or returned anything to the caller, so the builder has more flexibility in how it produces that Task. If you remember previously in the entry point method, the very last thing it does is return the Task to the caller, which it does by returning the result of accessing the builders Task property (so many things called “Task”, I know):
public Task CopyStreamToStreamAsync(Stream source, Stream destination){ ... return stateMachine.<>t__builder.Task;}
The builder knows if the method ever suspended, in which case it has a Task that was already created and just returns that. If the method never suspended and the builder doesnt yet have a task, it can manufacture a completed task here. In this case, with a successful completion, it can just use Task.CompletedTask rather than allocating a new task, avoiding any allocation. In the case of a generic Task<TResult>, the builder can just use Task.FromResult<TResult>(TResult result).
The builder can also do whatever translations it deems are appropriate to the kind of object its creating. For example, Task actually has three possible final states: success, failure, and canceled. The AsyncTaskMethodBuilders SetException method [special-cases](https://github.com/dotnet/runtime/blob/3e73be1b8082840545dbf85867cc4f9023e9b1aa/src/libraries/System.Private.CoreLib/src/System/Runtime/CompilerServices/AsyncTaskMethodBuilderT.cs#L461-L486) OperationCanceledException, transitioning the Task into a TaskStatus.Canceled final state if the exception provided is or derives from OperationCanceledException; otherwise, the task ends as TaskStatus.Faulted. Such a distinction often isnt apparent in consuming code; since the exception is stored into the Task regardless of whether its marked as Canceled or Faulted, code awaiting that Task will not be able to observe the difference between the states (the original exception will be propagated in either case)… it only affects code that interacts with the Task directly, such as via ContinueWith, which has overloads that enable a continuation to be invoked only for a subset of completion statuses.
Now that we understand the lifecycle aspects, heres everything filled in inside the try block in MoveNext:
private void MoveNext(){ try { int num = <>1__state; TaskAwaiter<int> awaiter; if (num != 0) { if (num != 1) { <buffer>5__2 = new byte[4096]; goto IL_008b; } awaiter = <>u__2; <>u__2 = default(TaskAwaiter<int>); num = (<>1__state = -1); goto IL_00f0; } TaskAwaiter awaiter2 = <>u__1; <>u__1 = default(TaskAwaiter); num = (<>1__state = -1); IL_0084: awaiter2.GetResult(); IL_008b: awaiter = source.ReadAsync(<buffer>5__2, 0, <buffer>5__2.Length).GetAwaiter(); if (!awaiter.IsCompleted) { num = (<>1__state = 1); <>u__2 = awaiter; <>t__builder.AwaitUnsafeOnCompleted(ref awaiter, ref this); return; } IL_00f0: int result; if ((result = awaiter.GetResult()) != 0) { awaiter2 = destination.WriteAsync(<buffer>5__2, 0, result).GetAwaiter(); if (!awaiter2.IsCompleted) { num = (<>1__state = 0); <>u__1 = awaiter2; <>t__builder.AwaitUnsafeOnCompleted(ref awaiter2, ref this); return; } goto IL_0084; } } catch (Exception exception) { <>1__state = -2; <buffer>5__2 = null; <>t__builder.SetException(exception); return; } <>1__state = -2; <buffer>5__2 = null; <>t__builder.SetResult();}
This kind of complication might feel a tad familiar. Remember how convoluted our manually-implemented BeginCopyStreamToStream based on APM was? This isnt quite as complicated, but is also way better in that the compiler is doing the work for us, having rewritten the method in a form of continuation passing while ensuring that all necessary state is preserved for those continuations. Even so, we can squint and follow along. Remember that the state was initialized to -1 in the entry point. We then enter MoveNext, find that this state (which is now stored in the num local) is neither 0 nor 1, and thus execute the code that creates the temporary buffer and then branches to label IL_008b, where it makes the call to stream.ReadAsync. Note that at this point were still running synchronously from this call to MoveNext, and thus synchronously from Start, and thus synchronously from the entry point, meaning the developers code called CopyStreamToStreamAsync and its still synchronously executing, having not yet returned back a Task to represent the eventual completion of this method. That might be about to change…
We call Stream.ReadAsync and we get back a Task<int> from it. The read may have completed synchronously, it may have completed asynchronously but so fast that its now already completed, or it might not have completed yet. Regardless, we have a Task<int> that represents its eventual completion, and the compiler emits code that inspects that Task<int> to determine how to proceed: if the Task<int> has in fact already completed (doesnt matter whether it was completed synchronously or just by the time we checked), then the code for this method can just continue running synchronously… no point in spending unnecessary overhead queueing a work item to handle the remainder of the methods execution when we can instead just keep running here and now. But to handle the case where the Task<int> hasnt completed, the compiler needs to emit code to hook up a continuation to the Task. It thus needs to emit code that asks the Task “are you done?” Does it talk to the Task directly to ask that?
It would be limiting if the only thing you could await in C# was a System.Threading.Tasks.Task. Similarly, it would be limiting if the C# compiler had to know about every possible type that could be awaited. Instead, C# does what it typically does in cases like this: it employs a pattern of APIs. Code can await anything that exposes that appropriate pattern, the “awaiter” pattern (just as you can foreach anything that provides the proper “enumerable” pattern). For example, we can augment the MyTask type we wrote earlier to implement the awaiter pattern:
class MyTask{ ... public MyTaskAwaiter GetAwaiter() => new MyTaskAwaiter { _task = this }; public struct MyTaskAwaiter : ICriticalNotifyCompletion { internal MyTask _task; public bool IsCompleted => _task._completed; public void OnCompleted(Action continuation) => _task.ContinueWith(_ => continuation()); public void UnsafeOnCompleted(Action continuation) => _task.ContinueWith(_ => continuation()); public void GetResult() => _task.Wait(); }}
A type can be awaited if it exposes a GetAwaiter() method, which Task does. That method needs to return something that in turn exposes several members, including an IsCompleted property, which is used to check at the moment IsCompleted is called whether the operation has already completed. And you can see that happening: at IL_008b, the Task returned from ReadAsync has GetAwaiter called on it, and then IsCompleted accessed on that struct awaiter instance. If IsCompleted returns true, then well end up falling through to IL_00f0, where the code calls another member of the awaiter: GetResult(). If the operation failed, GetResult() is responsible for throwing an exception in order to propagate it out of the await in the async method; otherwise, GetResult() is responsible for returning the result of the operation, if there is one. In the case here of the ReadAsync, if that result is 0, then we break out of our read/write loop, go to the end of the method where it calls SetResult, and were done.
Backing up a moment, though, the really interesting part of all of this is what happens if that IsCompleted check actually returns false. If it returns true, we just keep on processing the loop, akin to in the APM pattern when CompletedSynchronously returned true and the caller of the Begin method, rather than the callback, was responsible for continuing execution. But if IsCompleted returns false, we need to suspend the execution of the async method until the awaitd operation completes. That means returning out of MoveNext, and as this was part of Start and were still in the entry point method, that means returning the Task out to the caller. But before any of that can happen, we need to hook up a continuation to the Task being awaited (noting that to avoid stack dives as in the APM case, if the asynchronous operation completes after IsCompleted returns false but before we get to hook up the continuation, the continuation still needs to be invoked asynchronously from the calling thread, and thus itll get queued). Since we can await anything, we cant just talk to the Task instance directly; instead, we need to go through some pattern-based method to perform this.
Does that mean theres a method on the awaiter that will hook up the continuation? That would make sense; after all, Task itself supports continuations, has a ContinueWith method, etc… shouldnt it be the TaskAwaiter returned from GetAwaiter that exposes the method that lets us set up a continuation? It does, in fact. The awaiter pattern requires that the awaiter implement the INotifyCompletion interface, which contains a single method void OnCompleted(Action continuation). An awaiter can also optionally implement the ICriticalNotifyCompletion interface, which inherits INotifyCompletion and adds a void UnsafeOnCompleted(Action continuation) method. Per our previous discussion of ExecutionContext, you can guess what the difference between these two methods is: both hook up the continuation, but whereas OnCompleted should flow ExecutionContext, UnsafeOnCompleted neednt. The need for two distinct methods here, INotifyCompletion.OnCompleted and ICriticalNotifyCompletion.UnsafeOnCompleted, is largely historical, having to do with Code Access Security, or CAS. CAS no longer exists in .NET Core, and its off by default in .NET Framework, having teeth only if you opt back in to the legacy partial trust feature. When partial trust is used, CAS information flows as part of ExecutionContext, and thus not flowing it is “unsafe”, hence why methods that dont flow ExecutionContext were prefixed with “Unsafe”. Such methods were also attributed as [SecurityCritical], and partially trusted code cant call a [SecurityCritical] method. As a result, two variants of OnCompleted were created, with the compiler preferring to use UnsafeOnCompleted if provided, but with the OnCompleted variant always provided on its own in case an awaiter needed to support partial trust. From an async method perspective, however, the builder always flows ExecutionContext across await points, so an awaiter that also does so is unnecessary and duplicative work.
Ok, so the awaiter does expose a method to hook up the continuation. The compiler _could_ use it directly, except for a very critical piece of the puzzle: what exactly should the continuation be? And more to the point, with what object should it be associated? Remember that the state machine struct is on the stack, and the MoveNext invocation were currently running in is a method call on that instance. We need to preserve the state machine so that upon resumption we have all the correct state, which means the state machine cant just keep living on the stack; it needs to be copied to somewhere on the heap, since the stack is going to end up being used for other subsequent, unrelated work performed by this thread. And then the continuation needs to invoke the MoveNext method on that copy of the state machine on the heap.
Moreover, ExecutionContext is relevant here as well. The state machine needs to ensure that any ambient data stored in the ExecutionContext is captured at the point of suspension and then applied at the point of resumption, which means the continuation also needs to incorporate that ExecutionContext. So, just creating a delegate that points to MoveNext on the state machine is insufficient. Its also undesirable overhead. If when we suspend we create a delegate that points to MoveNext on the state machine, each time we do so well be boxing the state machine struct (even when its already on the heap as part of some other object) and allocating an additional delegate (the delegates this object reference will be to a newly boxed copy of the struct). We thus need to do a complicated dance whereby we ensure we only promote the struct from the stack to the heap the first time the method suspends execution but all other times uses the same heap object as the target of the MoveNext, and in the process ensures weve captured the right context, and upon resumption ensures were using that captured context to invoke the operation.
Thats a lot more logic than we want the compiler to emit… we instead want it encapsulated in a helper, for several reasons. First, its a lot of complicated code to be emitted into each users assembly. Second, we want to allow customization of that logic as part of implementing the builder pattern (well see an example of why later when talking about pooling). And third, we want to be able to evolve and improve that logic and have existing previously-compiled binaries just get better. Thats not a hypothetical; the library code for this support was completely overhauled in .NET Core 2.1, such that the operation is much more efficient than it was on .NET Framework. Well start by exploring exactly how this worked on .NET Framework, and then look at what happens now in .NET Core.
You can see in the code generated by the C# compiler happens when we need to suspend:
if (!awaiter.IsCompleted) // we need to suspend when IsCompleted is false{ <>1__state = 1; <>u__2 = awaiter; <>t__builder.AwaitUnsafeOnCompleted(ref awaiter, ref this); return;}
Were storing into the state field the state id that indicates the location we should jump to when the method resumes. Were then persisting the awaiter itself into a field, so that it can be used to call GetResult after resumption. And then just before returning out of the MoveNext call, the very last thing we do is call <>t__builder.AwaitUnsafeOnCompleted(ref awaiter, ref this), asking the builder to hook up a continuation to the awaiter for this state machine. (Note that it calls the builders AwaitUnsafeOnCompleted rather than the builders AwaitOnCompleted because the awaiter implements ICriticalNotifyCompletion; the state machine handles flowing ExecutionContext so we neednt require the awaiter to as well… as mentioned earlier, doing so would just be duplicative and unnecessary overhead.)
The implementation of that AwaitUnsafeOnCompleted method is too complicated to copy here, so Ill summarize [what it does](https://referencesource.microsoft.com/#mscorlib/system/runtime/compilerservices/AsyncMethodBuilder.cs,535) on .NET Framework:
1. It uses ExecutionContext.Capture() to grab the current context.
2. It then allocates a MoveNextRunner object to wrap both the captured context as well as the boxed state machine (which we dont yet have if this is the first time the method suspends, so we just use null as a placeholder).
3. It then creates an Action delegate to a Run method on that MoveNextRunner; this is how its able to get a delegate that will invoke the state machines MoveNext in the context of the captured ExecutionContext.
4. If this is the first time the method is suspending, we wont yet have a boxed state machine, so at this point it boxes it, creating a copy on the heap by storing the instance into a local typed as the IAsyncStateMachine interface. That box is then stored into the MoveNextRunner that was allocated.
5. Now comes a somewhat mind-bending step. If you look back at the definition of the state machine struct, it contains the builder, public AsyncTaskMethodBuilder <>t__builder;, and if you look at the definition of the builder, it contains internal IAsyncStateMachine m_stateMachine;. The builder needs to reference the boxed state machine so that on subsequent suspensions it can see its already boxed the state machine and doesnt need to do so again. But we just boxed the state machine, and that state machine contained a builder whose m_stateMachine field is null. We need to mutate that boxed state machines builders m_stateMachine to point to its parent box. To achieve that, the IAsyncStateMachine interface that the compiler-generated state machine struct implements includes a void SetStateMachine(IAsyncStateMachine stateMachine); method, and that state machine struct includes an implementation of that interface method:
private void SetStateMachine(IAsyncStateMachine stateMachine) => <>t__builder.SetStateMachine(stateMachine);
So the builder boxes the state machine, and then passes that box to the boxs SetStateMachine method, which calls to the builders SetStateMachine method, which stores the box into the field. Whew.
7. Finally, we have an Action that represents the continuation, and thats passed to the awaiters UnsafeOnCompleted method. In the case of a TaskAwaiter, the task will store that Action into the tasks continuation list, such that when the task completes, itll invoke the Action, call back through the MoveNextRunner.Run, call back through ExecutionContext.Run, and finally invoke the state machines MoveNext method to re-enter the state machine and continue running from where it left off.
Thats what happens on .NET Framework, and you can witness the outcome of this in a profiler, such as by running an allocation profiler to see whats allocated on each await. Lets take this silly program, which Ive written just to highlight the allocation costs involved:
using System.Threading;using System.Threading.Tasks;class Program{ static async Task Main() { var al = new AsyncLocal<int>() { Value = 42 }; for (int i = 0; i < 1000; i++) { await SomeMethodAsync(); } } static async Task SomeMethodAsync() { for (int i = 0; i < 1000; i++) { await Task.Yield(); } }}
This program is creating an AsyncLocal<int> to flow the value 42 through all subsequent async operations. Its then calling SomeMethodAsync 1000 times, each of which is suspending/resuming 1000 times. In Visual Studio, I run this using the [.NET Object Allocation Tracking profiler](https://learn.microsoft.com/visualstudio/profiling/dotnet-alloc-tool), which yields the following results:
[![Allocation associated with asynchronous operations on .NET Framework](Exported%20image%2020240808113926-2.png)](https://devblogs.microsoft.com/dotnet/wp-content/uploads/sites/10/2023/03/AllocationNetFramework.png)
Thats… a lot of allocation! Lets examine each of these to understand where theyre coming from.
- ExecutionContext. Theres over a million of these being allocated. Why? Because in .NET Framework, ExecutionContext is a _mutable_ data structure. Since we want to flow the data that was present at the time an async operation was forked and we dont want it to then see mutations performed after that fork, we need to copy the ExecutionContext. Every single forked operation requires such a copy, so with 1000 calls to SomeMethodAsync each of which is suspending/resuming 1000 times, we have a million ExecutionContext instances. Ouch.
- Action. Similarly, every time we await something thats not yet complete (which is the case with our million await Task.Yield()s), we end up allocating a new Action delegate to pass to that awaiters UnsafeOnCompleted method.
- MoveNextRunner. Same deal; theres a million of these, since in the outline of the steps earlier, every time we suspend, were allocating a new MoveNextRunner to store the Action and the ExecutionContext, in order to execute the former with the latter.
- LogicalCallContext. Another million. These are an implementation detail of AsyncLocal<T> on .NET Framework; AsyncLocal<T> stores its data into the ExecutionContexts “logical call context”, which is a fancy way of saying the general state thats flowed with the ExecutionContext. So, if were making a million copies of the ExecutionContext, were making a million copies of the LogicalCallContext, too.
- QueueUserWorkItemCallback. Each Task.Yield() is queueing a work item to the thread pool, resulting in a million allocations of the work item objects used to represent those million operations.
- Task<VoidResult>. Theres a thousand of these, so at least were out of the “million” club. Every async Task invocation that completes asynchronously needs to allocate a new Task instance to represent the eventual completion of that call.
- <SomeMethodAsync>d__1. This is the box of the compiler-generated state machine struct. 1000 methods suspend, 1000 boxes occur.
- QueueSegment/IThreadPoolWorkItem[]. There are several thousand of these, and theyre not technically related to async methods specifically, but rather to work being queued to the thread pool in general. In .NET Framework, the thread pools queue is a linked list of non-circular segments. These segments arent reused; for a segment of length N, once N work items have been enqueued into and dequeued from that segment, the segment is discarded and left up for garbage collection.
That was .NET Framework. [This](https://github.com/dotnet/runtime/blob/8de96c8b1b1cc3a781f23dcdf68c0aeb62dadbe7/src/libraries/System.Private.CoreLib/src/System/Runtime/CompilerServices/AsyncTaskMethodBuilderT.cs#L97-L145) is .NET Core:
[![Allocation associated with asynchronous operations on .NET Core](Exported%20image%2020240808113926-3.png)](https://devblogs.microsoft.com/dotnet/wp-content/uploads/sites/10/2023/03/AllocationNetCore.png)
So much prettier! For this sample on .NET Framework, there were more than 5 million allocations totaling ~145MB of allocated memory. For that same sample on .NET Core, there were instead only ~1000 allocations totaling only ~109KB. Why so much less?
- ExecutionContext. In .NET Core, ExecutionContext is now _immutable_. The downside to that is that every change to the context, e.g. by setting a value into an AsyncLocal<T>, requires allocating a new ExecutionContext. The upside, however, is that flowing context is way, way, way more common than is changing it, and as ExecutionContext is now immutable, we no longer need to clone as part of flowing it. “Capturing” the context is literally just reading it out of a field, rather than reading it and doing a clone of its contents. So its not only way, way, way more common to flow than to change, its also way, way, way cheaper.
- LogicalCallContext. This no longer exists in .NET Core. In .NET Core, the only thing ExecutionContext exists for is the storage for AsyncLocal<T>. Other things that had their own special place in ExecutionContext are modeled in terms of AsyncLocal<T>. For example, impersonation in .NET Framework would flow as part of the SecurityContext thats part of ExecutionContext; in .NET Core, impersonation flows via an AsyncLocal<SafeAccessTokenHandle> that uses a valueChangedHandler to make appropriate changes to the current thread.
- QueueSegment/IThreadPoolWorkItem[]. In .NET Core, the ThreadPools global queue is now implemented as a ConcurrentQueue<T>, and ConcurrentQueue<T> has been rewritten to be a linked list of _circular_ segments of non-fixed size. Once the size of a segment is large enough that the segment never fills because steady-state dequeues are able to keep up with steady-state enqueues, no additional segments need to be allocated, and the same large-enough segment is just used endlessly.
What about the rest of the allocations, like Action, MoveNextRunner, and <SomeMethodAsync>d__1? Understanding how the remaining allocations were removed requires diving into how this now works on .NET Core.
Lets rewind our discussion back to when we were discussing what happens at suspension time:
if (!awaiter.IsCompleted) // we need to suspend when IsCompleted is false{ <>1__state = 1; <>u__2 = awaiter; <>t__builder.AwaitUnsafeOnCompleted(ref awaiter, ref this); return;}
The code thats emitted here is the same regardless of which platform surface area is being targeted, so regardless of .NET Framework vs .NET Core, the generated IL for this suspension is identical. What changes, however, is the implementation of that AwaitUnsafeOnCompleted method, which on .NET Core is much different:
1. Things do start out the same: the method calls ExecutionContext.Capture() to get the current execution context.
2. Then things diverge from .NET Framework. The builder in .NET Core has just a single field on it:
public struct AsyncTaskMethodBuilder{ private Task<VoidTaskResult>? m_task; ...}
After capturing the ExecutionContext, it checks whether that m_task field contains an instance of an [AsyncStateMachineBox<TStateMachine>](https://github.com/dotnet/runtime/blob/8de96c8b1b1cc3a781f23dcdf68c0aeb62dadbe7/src/libraries/System.Private.CoreLib/src/System/Runtime/CompilerServices/AsyncTaskMethodBuilderT.cs#L273), where TStateMachine is the type of the compiler-generated state machine struct. That AsyncStateMachineBox<TStateMachine> type is the “magic.” Its defined like this:
private class AsyncStateMachineBox<TStateMachine> : Task<TResult>, IAsyncStateMachineBox where TStateMachine : IAsyncStateMachine{ private Action? _moveNextAction; public TStateMachine? StateMachine; public ExecutionContext? Context; ...}
Rather than having a separate Task, this _is_ the task (note its base type). Rather than boxing the state machine, the struct just lives as a strongly-typed field on this task. And rather than having a separate MoveNextRunner to store both the Action and the ExecutionContext, theyre just fields on this type, and since this _is_ the instance that gets stored into the builders m_task field, we have direct access to it and dont need to re-allocate things on every suspension. If the ExecutionContext changes, we can just overwrite the field with the new context and dont need to allocate anything else; any Action we have still points to the right place. So, after capturing the ExecutionContext, if we already have an instance of this AsyncStateMachineBox<TStateMachine>, this isnt the first time the method is suspending, and we can just store the newly captured ExecutionContext into it. If we dont already have an instance of AsyncStateMachineBox<TStateMachine>, then we need to allocate it:
var box = new AsyncStateMachineBox<TStateMachine>();taskField = box; // important: this must be done before storing stateMachine into box.StateMachine!box.StateMachine = stateMachine;box.Context = currentContext;
Note that line which the source comments as “important”. This takes the place of that complicated SetStateMachine dance in .NET Framework, such that SetStateMachine isnt actually used at all in .NET Core. The taskField you see there is a ref to the AsyncTaskMethodBuilders m_task field. We allocate the AsyncStateMachineBox<TStateMachine>, then via taskField store that object into the builders m_task (this is the builder thats in the state machine struct on the stack), and then copy that stack-based state machine (which now already contains the reference to the box) into the heap-based AsyncStateMachineBox<TStateMachine>, such that the AsyncStateMachineBox<TStateMachine> appropriately and recursively ends up referencing itself. Still mind bending, but a much more efficient mind bending.
4. We can then get an Action to a method on this instance that will invoke its MoveNext method that will do the appropriate ExecutionContext restoration prior to calling into the StateMachines MoveNext. And that Action can be cached into the _moveNextAction field such that any subsequent use can just reuse the same Action. That Action is then passed to the awaiters UnsafeOnCompleted to hook up the continuation.
That explanation explains why most of the rest of the allocations are gone: <SomeMethodAsync>d__1 doesnt get boxed and instead just lives as a field on the task itself, and the MoveNextRunner is no longer needed as it existed only to store the Action and ExecutionContext. But, based on this explanation, we should have still seen 1000 Action allocations, one per method call, and we didnt. Why? And what about those QueueUserWorkItemCallback objects… were still queueing as part of Task.Yield(), so why arent those showing up?
As I noted, one of the nice things about pushing off the implementation details into the core library is it can evolve the implementation over time, and weve already seen how it evolved from .NET Framework to .NET Core. Its also evolved further from the initial rewrite for .NET Core, with additional optimizations that benefit from having internal access to key components in the system. In particular, the async infrastructure knows about core types like Task and TaskAwaiter. And because it knows about them and has internals access, it doesnt have to play by the publicly-defined rules. The awaiter pattern followed by the C# language requires an awaiter to have an AwaitOnCompleted or AwaitUnsafeOnCompleted method, both of which take the continuation as an Action, and that means the infrastructure needs to be able to create an Action to represent the continuation, in order to work with arbitrary awaiters the infrastructure knows nothing about. But if the infrastructure encounters an awaiter it _does_ know about, its under no obligation to take the same code path. For all of the core awaiters defined in System.Private.CoreLib, then, the infrastructure has a leaner path it can follow, one that doesnt require an Action at all. These awaiters all know about IAsyncStateMachineBoxes, and are able to treat the box object itself as the continuation. So, for example, the YieldAwaitable returned by Task.Yield is able to queue the IAsyncStateMachineBox itself directly into the ThreadPool as a work item, and the TaskAwaiter used when awaiting a Task is able to store the IAsyncStateMachineBox itself directly into the Tasks continuation list. No Action needed, no QueueUserWorkItemCallback needed.
Thus, in the very common case where an async method only awaits things from System.Private.CoreLib (Task, Task<TResult>, ValueTask, ValueTask<TResult>, YieldAwaitable, and the ConfigureAwait variants of those), worst case is theres only ever a single allocation of overhead associated with the entire lifecycle of the async method: if the method ever suspends, it allocates that single Task-derived type which stores all other required state, and if the method never suspends, theres no additional allocation incurred.
We can get rid of that last allocation as well, if desired, at least in an amortized fashion. As has been shown, theres a default builder associated with Task (AsyncTaskMethodBuilder), and similarly theres a default builder associated with Task<TResult> (AsyncTaskMethodBuilder<TResult>) and with ValueTask and ValueTask<TResult> (AsyncValueTaskMethodBuilder and AsyncValueTaskMethodBuilder<TResult>, respectively). For ValueTask/ValueTask<TResult>, the builders are actually fairly simple, as they themselves only handle the synchronously-and-successfully-completing case, in which case the async method completes without ever suspending and the builders can just return a ValueTask.Completed or a ValueTask<TResult> wrapping the result value. For everything else, they just delegate to AsyncTaskMethodBuilder/AsyncTaskMethodBuilder<TResult>, since the ValueTask/ValueTask<TResult> thatll be returned just wraps a Task and it can share all of the same logic. But [.NET 6 and C# 10](https://devblogs.microsoft.com/dotnet/performance-improvements-in-net-6) introduced the ability for a method to override the builder thats used on a method-by-method basis, and introduced a couple of specialized builders for ValueTask/ValueTask<TResult> that are able to pool IValueTaskSource/IValueTaskSource<TResult> objects representing the eventual completion rather than using Tasks.
We can see the impact of this in our sample. Lets slightly tweak our SomeMethodAsync we were profiling to return ValueTask instead of Task:
static async ValueTask SomeMethodAsync(){ for (int i = 0; i < 1000; i++) { await Task.Yield(); }}
That will result in this generated entry point:
[AsyncStateMachine(typeof(<SomeMethodAsync>d__1))]private static ValueTask SomeMethodAsync(){ <SomeMethodAsync>d__1 stateMachine = default; stateMachine.<>t__builder = AsyncValueTaskMethodBuilder.Create(); stateMachine.<>1__state = -1; stateMachine.<>t__builder.Start(ref stateMachine); return stateMachine.<>t__builder.Task;}
Now, we add [AsyncMethodBuilder(typeof(PoolingAsyncValueTaskMethodBuilder))] to the declaration of SomeMethodAsync:
[AsyncMethodBuilder(typeof(PoolingAsyncValueTaskMethodBuilder))]static async ValueTask SomeMethodAsync(){ for (int i = 0; i < 1000; i++) { await Task.Yield(); }}
and the compiler instead outputs this:
[AsyncStateMachine(typeof(<SomeMethodAsync>d__1))][AsyncMethodBuilder(typeof(PoolingAsyncValueTaskMethodBuilder))]private static ValueTask SomeMethodAsync(){ <SomeMethodAsync>d__1 stateMachine = default; stateMachine.<>t__builder = PoolingAsyncValueTaskMethodBuilder.Create(); stateMachine.<>1__state = -1; stateMachine.<>t__builder.Start(ref stateMachine); return stateMachine.<>t__builder.Task;}
The actual C# code gen for the entirety of the implementation, including the whole state machine (not shown), is almost identical; the _only_ difference is the type of the builder thats created and stored and thus used everywhere we previously saw references to the builder. And if you look at [the code for](https://github.com/dotnet/runtime/blob/8de96c8b1b1cc3a781f23dcdf68c0aeb62dadbe7/src/libraries/System.Private.CoreLib/src/System/Runtime/CompilerServices/PoolingAsyncValueTaskMethodBuilderT.cs#L152-L218) PoolingAsyncValueTaskMethodBuilder, youll see its structure is almost identical to that of AsyncTaskMethodBuilder, including using some of the exact same shared routines for doing things like special-casing known awaiter types. The key difference is that instead of doing new AsyncStateMachineBox<TStateMachine>() when the method first suspends, it instead does StateMachineBox<TStateMachine>.RentFromCache(), and upon the async method (SomeMethodAsync) completing and an await on the returned ValueTask completing, the rented box is returned to the cache. That means (amortized) zero allocation:
[![Allocation associated with asynchronous operations on .NET Core with pooling](Exported%20image%2020240808113926-4.png)](https://devblogs.microsoft.com/dotnet/wp-content/uploads/sites/10/2023/03/AllocationNetCoreWithPooling.png)
That cache in and of itself is a bit interesting. Object pooling can be a good idea and it can be a bad idea. The more expensive an object is to create, the more valuable it is to pool them; so, for example, its a lot more valuable to pool really large arrays than it is to pool really tiny arrays, because larger arrays not only require more CPU cycles and memory accesses to zero out, they put more pressure on the garbage collector to collect more often. For very small objects, though, pooling them can be a net negative. Pools are just memory allocators, as is the GC, so when you pool, youre trading off the costs associated with one allocator for the costs associated with another, and the GC is very efficient at handling lots of tiny, short-lived objects. If you do a lot of work in an objects constructor, avoiding that work can dwarf the costs of the allocator itself, making pooling valuable. But if you do little to no work in an objects constructor, and you pool it, youre betting that your allocator (your pool) is more efficient for the access patterns employed than is the GC, and that is frequently a bad bet. There are other costs involved as well, and in some cases you can end up effectively fighting against the GCs heuristics; for example, the GC is optimized based on the premise that references from higher generation (e.g. gen2) objects to lower generation (e.g. gen0) objects are relatively rare, but pooling objects can invalidate those premises.
Now, the objects created by async methods arent _tiny_, and they can be on super hot paths, so pooling can be reasonable. But to make it as valuable as possible we also want to avoid as much overhead as possible. The pool is thus very simple, opting to make renting and returning really fast with little to no contention, even if that means it might end up allocating more than it would if it more aggressively cached more. For each state machine type, the implementation [pools](https://github.com/dotnet/runtime/blob/8de96c8b1b1cc3a781f23dcdf68c0aeb62dadbe7/src/libraries/System.Private.CoreLib/src/System/Runtime/CompilerServices/PoolingAsyncValueTaskMethodBuilderT.cs#L287-L292) up to a single state machine box per _thread_ and a single state machine box per _core_; this allows it to rent and return with minimal overhead and minimal contention (no other thread can be accessing the thread-specific cache at the same time, and its rare for another thread to be accessing the core-specific cache at the same time). And while this might seem like a relatively small pool, its also quite effective at significantly reducing steady state allocation, given that the pool is only responsible for storing objects not currently in use; you could have a million async methods all in flight at any given time, and even though the pool is only able to store up to one object per thread and per core, it can still avoid dropping lots of objects, since it only needs to store an object long enough to transfer it from one operation to another, not while its in use by that operation.
### SynchronizationContext and ConfigureAwait
We talked about SynchronizationContext previously in the context of the EAP pattern and mentioned that it would show up again. SynchronizationContext makes it possible to call reusable helpers and automatically be scheduled back whenever and to wherever the calling environment deems fit. As a result, its natural to expect that to “just work” with async/await, and it does. Going back to our button click handler from earlier:
ThreadPool.QueueUserWorkItem(_ =>{ string message = ComputeMessage(); button1.BeginInvoke(() => { button1.Text = message; });});
with async/await wed like to instead be able to write this as follows:
button1.Text = await Task.Run(() => ComputeMessage());
That invocation of ComputeMessage is offloaded to the thread pool, and upon the methods completion, execution transitions back to the UI thread associated with the button, and the setting of its Text property happens on that thread.
That integration with SynchronizationContext is left up to the awaiter implementation (the code generated for the state machine knows nothing about SynchronizationContext), as its the awaiter that is responsible for actually invoking or queueing the supplied continuation when the represented asynchronous operation completes. While a custom awaiter need not respect SynchronizationContext.Current, the awaiters for Task, Task<TResult>, ValueTask, and ValueTask<TResult> all do. That means that, by default, when you await a Task, a Task<TResult>, a ValueTask, a ValueTask<TResult>, or even the result of a Task.Yield() call, the awaiter by default will look up the current SynchronizationContext and then if it successfully got a non-default one, will eventually queue the continuation to that context.
We can see this if we look at the code involved in TaskAwaiter. Heres a snippet of the [relevant code](https://github.com/dotnet/runtime/blob/967a59712996c2cdb8ce2f65fb3167afbd8b01f3/src/libraries/System.Private.CoreLib/src/System/Threading/Tasks/Task.cs#L2558-L2583) from Corelib:
internal void UnsafeSetContinuationForAwait(IAsyncStateMachineBox stateMachineBox, bool continueOnCapturedContext){ if (continueOnCapturedContext) { SynchronizationContext? syncCtx = SynchronizationContext.Current; if (syncCtx != null && syncCtx.GetType() != typeof(SynchronizationContext)) { var tc = new SynchronizationContextAwaitTaskContinuation(syncCtx, stateMachineBox.MoveNextAction, flowExecutionContext: false); if (!AddTaskContinuation(tc, addBeforeOthers: false)) { tc.Run(this, canInlineContinuationTask: false); } return; } else { TaskScheduler? scheduler = TaskScheduler.InternalCurrent; if (scheduler != null && scheduler != TaskScheduler.Default) { var tc = new TaskSchedulerAwaitTaskContinuation(scheduler, stateMachineBox.MoveNextAction, flowExecutionContext: false); if (!AddTaskContinuation(tc, addBeforeOthers: false)) { tc.Run(this, canInlineContinuationTask: false); } return; } } } ...}
This is part of a method thats determining what object to store into the Task as a continuation. Its being passed the stateMachineBox, which, as was alluded to earlier, can be stored directly into the Tasks continuation list. However, this special logic might wrap that IAsyncStateMachineBox to also incorporate a scheduler if one is present. It checks to see whether theres currently a non-default SynchronizationContext, and if there is, it creates a SynchronizationContextAwaitTaskContinuation as the actual object thatll be stored as the continuation; that object in turn wraps the original and the captured SynchronizationContext, and knows how to invoke the formers MoveNext in a work item queued to the latter. This is how youre able to await as part of some event handler in a UI application and have the code after the awaits completion continue on the right thread. The next interesting thing to note here is that its not just paying attention to a SynchronizationContext: if it couldnt find a custom SynchronizationContext to use, it also looks to see whether the TaskScheduler type thats used by Tasks has a custom one in play that needs to be considered. As with SynchronizationContext, if theres a non-default one of those, its then wrapped with the original box in a TaskSchedulerAwaitTaskContinuation thats used as the continuation object.
But arguably the most interesting thing to notice here is the very first line of the method body: if (continueOnCapturedContext). We only do these checks for SynchronizationContext/TaskScheduler if continueOnCapturedContext is true; if its false, the implementation behaves as if both were default and ignores them. What, pray tell, sets continueOnCapturedContext to false? Youve probably guessed it: using the ever popular ConfigureAwait(false).
I talk about ConfigureAwait at length in [ConfigureAwait FAQ](https://devblogs.microsoft.com/dotnet/configureawait-faq/), so Id encourage you to read that for more information. Suffice it to say, the _only_ thing ConfigureAwait(false) does as part of an await is feed its argument Boolean into this function (and others like it) as that continueOnCapturedContext value, so as to skip the checks on SynchronizationContext/TaskScheduler and behave as if neither of them existed. In the case of Tasks, this then permits the Task to invoke its continuations wherever it deems fit rather than being forced to queue them to execute on some specific scheduler.
I previously mentioned one other aspect of SynchronizationContext, and I said wed see it again: OperationStarted/OperationCompleted. Nows the time. These rear their heads as part of the feature everyone loves to hate: async void. ConfigureAwait-aside, async void is arguably one of the most divisive features added as part of async/await. It was added for one reason and one reason only: event handlers. In a UI application, you want to be able to write code like the following:
button1.Click += async (sender, eventArgs) =>{ button1.Text = await Task.Run(() => ComputeMessage()); };
but if all async methods had to have a return type like Task, you wouldnt be able to do this. The Click event has a signature public event EventHandler? Click;, with EventHandler defined as public delegate void EventHandler(object? sender, EventArgs e);, and thus to provide a method that matches that signature, the method needs to be void-returning.
There are a variety of reasons async void is considered bad, why [articles](https://learn.microsoft.com/archive/msdn-magazine/2013/march/async-await-best-practices-in-asynchronous-programming) recommend avoiding it wherever possible, and why [analyzers](https://github.com/microsoft/vs-threading/blob/main/doc/analyzers/VSTHRD101.md) have sprung up to flag use of them. One of the biggest issues is with delegate inference. Consider this program:
using System.Diagnostics;Time(async () =>{ Console.WriteLine("Enter"); await Task.Delay(TimeSpan.FromSeconds(10)); Console.WriteLine("Exit");});static void Time(Action action){ Console.WriteLine("Timing..."); Stopwatch sw = Stopwatch.StartNew(); action(); Console.WriteLine($"...done timing: {sw.Elapsed}");}
One could easily expect this to output an elapsed time of at least 10 seconds, but if you run this youll instead find output like this:
Timing...Enter...done timing: 00:00:00.0037550
Huh? Of course, based on everything weve discussed in this post, it should be understood what the problem is. The async lambda is actually an async void method. Async methods return to their caller the moment they hit the first suspension point. If this were an async Task method, thats when the Task would be returned. But in the case of an async void, nothing is returned. All the Time method knows is that it invoked action(); and the delegate call returned; it has no idea that the async method is actually still “running” and will asynchronously complete later.
Thats where OperationStarted/OperationCompleted come in. Such async void methods are similar in nature to the EAP methods discussed earlier: the initiation of such methods is void, and so you need some other mechanism to be able to track all such operations in flight. The EAP implementations thus call the current SynchronizationContexts OperationStarted when the operation is initiated and OperationCompleted when it completes, and async void does the same. The builder associated with async void is AsyncVoidMethodBuilder. Remember in the entry point of an async method how the compiler-generated code invokes the builders static Create method to get an appropriate builder instance? AsyncVoidMethodBuilder takes advantage of that in order to hook creation and invoke OperationStarted:
public static AsyncVoidMethodBuilder Create(){ SynchronizationContext? sc = SynchronizationContext.Current; sc?.OperationStarted(); return new AsyncVoidMethodBuilder() { _synchronizationContext = sc };}
Similarly, when the builder is marked for completion via either SetResult or SetException, it invokes the corresponding OperationCompleted method. This is how a unit testing framework like xunit is able to have async void test methods and still employ a maximum degree of concurrency on concurrent test executions, for example in xunits [AsyncTestSyncContext](https://github.com/xunit/xunit/blob/4d1f2e5d4ac9260487d0a8f35a2d045388021b33/src/xunit.v3.core/Sdk/AsyncTestSyncContext.cs#L1).
With that knowledge, we can now rewrite our timing sample:
using System.Diagnostics;Time(async () =>{ Console.WriteLine("Enter"); await Task.Delay(TimeSpan.FromSeconds(10)); Console.WriteLine("Exit");});static void Time(Action action){ var oldCtx = SynchronizationContext.Current; try { var newCtx = new CountdownContext(); SynchronizationContext.SetSynchronizationContext(newCtx); Console.WriteLine("Timing..."); Stopwatch sw = Stopwatch.StartNew(); action(); newCtx.SignalAndWait(); Console.WriteLine($"...done timing: {sw.Elapsed}"); } finally { SynchronizationContext.SetSynchronizationContext(oldCtx); }}sealed class CountdownContext : SynchronizationContext{ private readonly ManualResetEventSlim _mres = new ManualResetEventSlim(false); private int _remaining = 1; public override void OperationStarted() => Interlocked.Increment(ref _remaining); public override void OperationCompleted() { if (Interlocked.Decrement(ref _remaining) == 0) { _mres.Set(); } } public void SignalAndWait() { OperationCompleted(); _mres.Wait(); }}
Here, Ive created a SynchronizationContext that tracks a count for pending operations, and supports blocking waiting for them all to complete. When I run that, I get output like this:
Timing...EnterExit...done timing: 00:00:10.0149074
Tada!
### State Machine Fields
At this point, weve seen the generated entry point method and how everything in the MoveNext implementation works. We also glimpsed some of the fields defined on the state machine. Lets take a closer look at those.
For the CopyStreamToStream method shown earlier:
public async Task CopyStreamToStreamAsync(Stream source, Stream destination){ var buffer = new byte[0x1000]; int numRead; while ((numRead = await source.ReadAsync(buffer, 0, buffer.Length)) != 0) { await destination.WriteAsync(buffer, 0, numRead); }}
here are the fields we ended up with:
private struct <CopyStreamToStreamAsync>d__0 : IAsyncStateMachine{ public int <>1__state; public AsyncTaskMethodBuilder <>t__builder; public Stream source; public Stream destination; private byte[] <buffer>5__2; private TaskAwaiter <>u__1; private TaskAwaiter<int> <>u__2; ...}
What are each of these?
- <>1__state. The is the “state” in “state machine”. It defines the current state the state machine is in, and most importantly what should be done the next time MoveNext is called. If the state is -2, the operation has completed. If the state is -1, either were about to call MoveNext for the first time or MoveNext code is currently running on some thread. If youre debugging an async methods processing and you see the state as -1, that means theres some thread somewhere thats actually executing the code contained in the method. If the state is 0 or greater, the method is suspended, and the value of the state tells you at which await its suspended. While this isnt a hard and fast rule (certain code patterns can confuse the numbering), in general the state assigned corresponds to the 0-based number of the await in top-to-bottom ordering of the source code. So, for example, if the body of an async method were entirely: await A();await B();await C();await D();
and you found the state value was 2, that almost certainly means the async method is currently suspended waiting for the task returned from C() to complete.
- <>t__builder. This is the builder for the state machine, e.g. AsyncTaskMethodBuilder for a Task, AsyncValueTaskMethodBuilder<TResult> for a ValueTask<TResult>, AsyncVoidMethodBuilder for an async void method, or whatever builder was declared for use via [AsyncMethodBuilder(...)] on either the async return type or overridden via such an attribute on the async method itself. As previously discussed, the builder is responsible for the lifecycle of the async method, including creating the return task, eventually completing that task, and serving as an intermediary for suspension, with the code in the async method asking the builder to suspend until a specific awaiter completes.
- source/destination. These are the method parameters. You can tell because theyre not name mangled; the compiler has named them exactly as the parameter names were specified. As noted earlier, all parameters that are used by the method body need to be stored onto the state machine so that the MoveNext method has access to them. Note I said “used by”. If the compiler sees that a parameter is unused by the body of the async method, it can optimize away the need to store the field. For example, given the method: public async Task M(int someArgument){ await Task.Yield();}
the compiler will emit these fields onto the state machine:
private struct <M>d__0 : IAsyncStateMachine{ public int <>1__state; public AsyncTaskMethodBuilder <>t__builder; private YieldAwaitable.YieldAwaiter <>u__1; ...}
Note the distinct lack of something named someArgument. But, if we change the async method to actually use the argument in any way:
public async Task M(int someArgument){ Console.WriteLine(someArgument); await Task.Yield();}
it shows up:
private struct <M>d__0 : IAsyncStateMachine{ public int <>1__state; public AsyncTaskMethodBuilder <>t__builder; public int someArgument; private YieldAwaitable.YieldAwaiter <>u__1; ...}
- <buffer>5__2;. This is the buffer “local” that got lifted to be a field so that it could survive across await points. The compiler tries reasonably hard to keep state from being lifted unnecessarily. Note that theres another local in the source, numRead, that _doesnt_ have a corresponding field in the state machine. Why? Because its not necessary. That local is set as the result of the ReadAsync call and is then used as the input to the WriteAsync call. Theres no await in between those and across which the numRead value would need to be stored. Just as how in a synchronous method the JIT compiler could choose to store such a value entirely in a register and never actually spill it to the stack, the C# compiler can avoid lifting this local to be a field as it neednt preserve its value across any awaits. In general, the C# compiler can elide lifting locals if it can prove that their value neednt be preserved across awaits.
- <>u__1 and <>u__2. There are two awaits in the async method: one for a Task<int> returned by ReadAsync, and one for a Task returned by WriteAsync. Task.GetAwaiter() returns a TaskAwaiter, and Task<TResult>.GetAwaiter() returns a TaskAwaiter<TResult>, both of which are distinct struct types. Since the compiler needs to get these awaiters prior to the await (IsCompleted, UnsafeOnCompleted) and then needs to access them after the await (GetResult), the awaiters need to be stored . And since theyre distinct struct types, the compiler needs to maintain two separate fields to do so (the alternative would be to box them and have a single object field for awaiters, but that would result in extra allocation costs). The compiler will try to reuse fields whenever possible, though. If I have: public async Task M(){ await Task.FromResult(1); await Task.FromResult(true); await Task.FromResult(2); await Task.FromResult(false); await Task.FromResult(3);}
there are five awaits, but only two different types of awaiters involved: three are TaskAwaiter<int> and two are TaskAwaiter<bool>. As such, there only end up being two awaiter fields on the state machine:
private struct <M>d__0 : IAsyncStateMachine{ public int <>1__state; public AsyncTaskMethodBuilder <>t__builder; private TaskAwaiter<int> <>u__1; private TaskAwaiter<bool> <>u__2; ...}
Then if I change my example to instead be:
public async Task M(){ await Task.FromResult(1); await Task.FromResult(true); await Task.FromResult(2).ConfigureAwait(false); await Task.FromResult(false).ConfigureAwait(false); await Task.FromResult(3);}
there are still only Task<int>s and Task<bool>s involved, but Im actually using four distinct struct awaiter types, because the awaiter returned from the GetAwaiter() call on the thing returned by ConfigureAwait is a different type than that returned by Task.GetAwaiter()… this is again evident from the awaiter fields created by the compiler:
private struct <M>d__0 : IAsyncStateMachine{ public int <>1__state; public AsyncTaskMethodBuilder <>t__builder; private TaskAwaiter<int> <>u__1; private TaskAwaiter<bool> <>u__2; private ConfiguredTaskAwaitable<int>.ConfiguredTaskAwaiter <>u__3; private ConfiguredTaskAwaitable<bool>.ConfiguredTaskAwaiter <>u__4; ...}
If you find yourself wanting to optimize the size associated with an async state machine, one thing you can look at is whether you can consolidate the kinds of things being awaited and thereby consolidate these awaiter fields.
There are other kinds of fields you might see defined on a state machine. Notably, you might see some fields containing the word “wrap”. Consider this silly example:
public async Task<int> M() => await Task.FromResult(42) + DateTime.Now.Second;
This produces a state machine with the following fields:
private struct <M>d__0 : IAsyncStateMachine{ public int <>1__state; public AsyncTaskMethodBuilder<int> <>t__builder; private TaskAwaiter<int> <>u__1; ...}
Nothing special so far. Now flip the order of the expressions being added:
public async Task<int> M() => DateTime.Now.Second + await Task.FromResult(42);
With that, you get these fields:
private struct <M>d__0 : IAsyncStateMachine{ public int <>1__state; public AsyncTaskMethodBuilder<int> <>t__builder; private int <>7__wrap1; private TaskAwaiter<int> <>u__1; ...}
We now have one more: <>7__wrap1. Why? Because we computed the value of DateTime.Now.Second, and only after computing it, we had to await something, and the value of the first expression needs to be preserved in order to add it to the result of the second. The compiler thus needs to ensure that the temporary result from that first expression is available to add to the result of the await, which means it needs to spill the result of the expression into a temporary, which it does with this <>7__wrap1 field. If you ever find yourself hyper-optimizing async method implementations to drive down the amount of memory allocated, you can look for such fields and see if small tweaks to the source could avoid the need for spilling and thus avoid the need for such temporaries.
## Wrap Up
I hope this post has helped to illuminate exactly whats going on under the covers when you use async/await, but thankfully you generally dont need to know or care. There are many moving pieces here, all coming together to create an efficient solution to writing scalable asynchronous code without having to deal with callback soup. And yet at the end of the day, those pieces are actually relatively simple: a universal representation for any asynchronous operation, a language and compiler capable of rewriting normal control flow into a state machine implementation of coroutines, and patterns that bind them all together. Everything else is optimization gravy.
Happy coding!

View File

@@ -0,0 +1,52 @@
Clipped from: [https://ferd.ca/notes/paper-moving-off-the-map.html](https://ferd.ca/notes/paper-moving-off-the-map.html)
These are notes I have taken elsewhere that I'm re-posting here as my weekly paper reading because it is quite simply one of my favourite papers ever.
[Moving off the Map: How Knowledge of Organizational Operations Empowers and Alienates](https://sci-hub.se/10.1287/orsc.2018.1277) is a work of ethnography where a researcher embedded herself into 5 organizations including 6 projects aiming at re-structuration—business process redesign (BRP)—and which all had a phase of extensive process mapping. She noticed that at the projects' conclusions, most employees returned to their roles and got raises within the organization, but a subset of them who were centrally located within the organization decided to move to peripheral roles. She decided to investigate this.
What she found was that tracing out the structure of how work is done (and what work is done) and decisions are made was a significant activity behind the split. It happened because people doing this tracing activity had a shock when they realized that the business' structure was not coordinated nor planned, but an emergent mess and consequence of local behaviours in various groups. Their new understanding of work resulted in either Empowerment ("I now know how I can change things") or Alienation ("Nothing I thought mattered does, my work is useless here"), which explained their move to peripheral roles.
Some of these reports are also just plain heart breaking. I have so many highlights for it.
The paper starts by mentioning that centrally-located actors (people at the core of the management structure of an organization) are less likely to initiate change, and more likely to stall it. Additionally, desire for change are likely to come from the periphery, and as people move towards the center, that desire tends to go away. This is a surprise to no one.
However, if central actors are to initiate change, it comes from either a) contradictions, tensions, or inconsistencies are experienced and push them to reflections, and b) being exposed to how other organizations (or even societies) do things, which opens more awareness. These two things are called "disembedding", and can lead to central actors pushing for structural change.
The paper accidentally "discovered" a third approach: taking the time to study how things are done in the organization can cause that dissonance, and encourage central actors to move to the periphery of the organization _in order to effect change_ because they lose trust in the structure of the organization itself and their role in it.
This was found out while the author was doing a study of 5 big corporations with 6 major business restructuration projects with hundreds of workers, and she noticed that while some employees went back to their roles (but with promotions), or towards roles that were more central when it was done, a subset of employees instead left very central roles to go work on the periphery, for sometimes less interesting conditions. So she started asking why and ran a big analysis.
What she noticed is that all the employees who eventually left their roles were assigned different specific tasks from the rest of people in these projects: they had been ask to do process mapping, where essentially they had to make a representation of "what we do here", how the business works, how decisions are made, and how information moves around. People not involved didn't find it significant, but people involved were shocked into leaving their roles, to make it short.
The author makes a point that it's not process mapping causing this, but rather that having a deep engagement in representing and understanding the operations of the organization and how their own role would fit in it would cause this to happen—it was probabilistic.
The tracing was done by employees who would do things like walk the floor, ask people how they do work, sit in meetings with question like "What do we do?" with people in various roles, asked them to list tasks on whiteboards, connecting them with strings, and consolidated into huge maps like the following, which connected local experiences into a broader organizational context:
![Walls covered in colorful sheet of papers with strings between all elements. This is an implementation of process-mapping out the organization and who covers what](Exported%20image%2020240808113929-0.png)
This had the effect of surface things that were previously invisible and make it discrete. This likely ties into concepts mentioned here before of "work as done" vs. "work as imagined":
The map allowed them to see how the system operated below the surface, integrating all the pieces to generate a comprehensive view. They commented on the uniqueness of this comprehensive view: “We dont allow people to see the end-to-end view... to see how things interrelate.” One explained that the experience “ruins [ones] perspective in a good way.” Another described how it gave her a “whole different way of looking at things.” By revealing the web of roles, relations, and routines that coalesce to make the organization, the map made the organizations actual operation intelligible.
[...]
Competent members of organizations draw on everyday knowledge [...] as they perform their roles, but this knowledge does not speak to the organizations broader order. Despite how remarkably capable these employees were at “recognizing, knowing, and doing the lived order,” the broader order or structure is often “resistance to analytic recovery” from the inside. Even if they would like to observe and reflect on their organizations detailed operating process, they rarely have opportunities, such as building process maps, that provide time and access.
So what were the immediate consequences? I'm quoting this directly:
They expected to observe inefficiencies and waste, the targets of redesign, and they did. Tasks that could be done with one or two hand-offs were taking three or four. Data painstakingly collected for decision-making processes were not used. Local repairs to work processes in one unit were causing downstream problems in another. Workarounds, duplication of effort, and poor communication and coordination were all evident on the map.
Beyond these issues, they observed a more fundamental problem. A team member explained, “Im getting a really clear visual of what the mess is.” Standing back from the wall, he sighed, and said, “The problem is that it was not designed in the first place.” Instead of observing a system designed, adapted, and coordinated to achieve stated goals, he pointed to three examples on the map that demonstrated the exercise of agency in various places and at various levels in the organization. These change efforts lacked broader perspective and direction as well as coordination and integration with other efforts
They mention examples such as a "kingdom builder" where the map revealed some manager who kept accumulating departments for the sake of accumulating power but was invisible to the organization, and essentially just found a lot of "what the fuck, this is just random shit that's leftovers from really old decisions." People see local problems, general approaches, and they try to fix things. This clashes with things the organization tries to do (when it tries), and there is no coherent organization to anything:
Some held out hope that one or two people at the top knew of these design and operation issues; however, they were often disabused of this optimism. For example, a manager walked the CEO through the map, presenting him with a view he had never seen before and illustrating for him the lack of design and the disconnect between strategy and operations. **The CEO, after being walked through the map, sat down, put his head on the table, and said, “This is even more fucked up than I imagined.”** The CEO revealed that not only was the operation of his organization out of his control but that his grasp on it was imaginary.
They learned that what they had previously attributed to the direction and control of centralized, bureaucratic forces was actually the aggregation of the work and decisions of people distributed throughout the organization. Everyone was working on the part of the organization that they were familiar with, assuming that another set of people were attending to the larger picture, coordinating the larger system to achieve goals and keeping the organization operating. They found out that this was not the case.
This may not necessarily be surprising to people, but it may be surprising for people to learn that CEOs and others think they have so much more control than they do!
Anyway, the two reactions in general were either Empowerment or Alienation.
On the front of Empowerment, this is caused because:
Members of the organization carry on as though these distinctions are facts, burdening the organizations categories, practices, and boundaries with a false sense of durability and purpose.
[...]
The idea that organizations are an ongoing human product was a provocative insight for these employees. This new perspective, as one explained, “made things seem possible.” Once they could see the “what” as a dynamic social creation, they could begin asking better questions about “how.” A team member explained that the logic of organization should not be fixed and how its rules, synthetic creations, are free to deviate
[...]
Their peripheral role choices allowed team members to exploit this new understanding of the organizations operations. They could work with new assumptions about the mutability and possibility of the organization and create structures and systems to coordinate and direct the web of roles and interactions. Their new role choices also allowed them to remain above and outside of the organizations daily operations.
So in short, getting how a lot of it isn't fixed, how a lot of it is arbitrary but flexible meant that these people felt they understood how to effect change better, and that by moving away from the center and into the periphery, they could start doing effective change work.
Alienation is so god damn heartbreaking though, and the author warns that before starting this process in an organization, you have to be ready that some people may feel a major shock that the work they thought was valuable and important is in fact useless and worth nothing. In fact the author warns that finding work and jobs that were not meaningful or useful at all was a common theme:
As part of the map-building process, employees were invited to identify their role on the map and to indicate how it was connected to other roles through either inputs or outputs. Team members recounted that it was difficult to observe employees “go through a real emotional struggle when they see that what they are doing is not really adding value or that what they are doing is really disconnected from what they thought they were doing.” In one case, a finance manager noticed that his role was on the wall but that it was not connected to any other role on the wall. He had been producing financial reports and sending them to several departments because he understood them to be crucial for their decision-making process; however, no one had identified his work as an input to theirs.
This realization was, in the end, **devastating for him. He was on the verge of tears... at first, he became very argumentative and was trying to convince people that you go from this Post-it note down here to mine. [Other employees explained] Well no, we dont do that. It was a two-hour conversation. And he finally sat down, and he said, so why am I doing this? It was devastating.**
The eventual outcome of this “aha” was that the manager was moved to another role in the department after working four years in a position that had served almost no purpose.
After such analyses, team members could not look at particular roles and people in the same way.
A lot of people also found out that they thought they were solving real problems, helping people with real issues, finding real work-arounds, but found that in the overall organizational map, it was meaningless and had no impact: they could be fixing real problems in departments that themselves were not useful.
Others found that they had properly fixed issues by introducing new databases with critical information, but that they had been unable to get any buy-in for that, so analysts and people having spent a lot of time on these just had no impact at all:
Their knowledge of the limits of local, small-scale change and the futility of changing parts of the organization without addressing the system as a whole, discouraged employees from returning to their career in the organization. They did not want to contribute to the mess or reproduce the mess they had observed.
[...]
What they had learned could not be unlearned or ignored.
The author state that whether it is due to alienation or empowerment, both behaviours push people to move to the edges of the system, where they can either find new roles or types of changes that they believe are more useful. The structural knowledge gain essentially let them know of better ways to do useful things and enact change. Specifically, learning that the organization's structure is the result of interactions rather than a context in which they take place is a key learning that sociologists knew already:
This perspective or comprehension affects how we speak and act. We speak about organizations as if they are objects that exist independent of us, and we act as though they constrain and guide our actions. When we objectify social systems (organizations, communities, families, gender roles), we apprehend them as “prearranged patterns” that impose themselves on us, coercing particular roles and rules. We free ourselves to talk about and inhabit them as independent of us: as existing prior to us, standing before us, outliving us, and operating without us. Given this, we are relieved of greater responsibility for them. Our responsibility is to skillfully fulfill our role within these objectified realms.
[...]
Whereas, as some sociologists “know that organizations and institutions exist only in actual peoples doings and that these are necessarily particular, local and ephemeral”, employees may be less likely to know this. When they do, it problematizes their past and future participation.
[...]
The realization that social worlds do not have an independent, stable existence but instead emerge from our collective action is “sometimes arrived at in a moment of heady delight, but often as a horrifying realization”. This realization is considered a “fatal insight” because it destroys assumptions that the current order, roles, rules, and routines are given. Within the system of roles, rules, and routines, there is far more room to maneuver than previously assumed. Rejection of objectivity puts possibility, perhaps even responsibility, squarely in the court of subjectivity.
I think this quote above is real good.
I'm going to conclude with it (, although the author adds a bit of a section about mentioning that given this research means that we can suspect some of the most effective change to be driven by actors who once were at the core of the system and moved to its periphery. This likely is a sign that they know how shit works and have an idea of how to challenge it. Insider knowledge dragged to the edges may be a key for strong means to modifying how things work. I'll let you read the paper if you want the details of that.

View File

@@ -0,0 +1,58 @@
Clipped from: [https://lethain.com/intro-product-management/](https://lethain.com/intro-product-management/)
![Problem Discovery Problem Selection Problem Discovery Solution Execution Validation Problem Selection Solution Execution Validation Problem Discovery Problem Selection Problem Discoverv Solution Execution Validation Problem Selection Solution Ex Validation ](Exported%20image%2020240808113923-0.png)
Most engineering organizations separate engineering and product leadership into distinct roles. This is usually ideal, not only because these roles benefit on distinct skills, but also because they thrive from different perspectives and priorities. Its quite hard to do both well at the same time.
Ive met many product managers who are excellent operators, but few product managers who can operate at a high degree while also getting deep with their users needs. Likewise, Ive worked with many engineering managers who ground their work in their users needs, but few who can affix their attention on those users when things start getting rocky within their team.
Reality isnt always accommodating of this ideal setup. Maybe your teams product manager leaves or a [new team is being formed](https://lethain.com/durably-excellent-teams/), and you, as an engineering leader, need to cover both roles for a few months. This can be exciting, and yes, this can be a time when “exciting” rhymes with “terrifying.”
Product management is a deep profession, and mastery requires years of practice, but Ive developed a simple framework for product management to use when Ive found [myself fulfilling product management](https://lethain.com/product-management-infra-engineering/) responsibilities for a team. Its not perfect, but hopefully itll be useful for you as well.
Product management is an iterative elimination tournament, with each round consisting of _problem discovery_, _problem selection_ and _solution validation_. _Problem discovery_ is uncovering possible problems to work on, _problem selection_ is filtering those problems down to a viable subset, and _solution validation_ is ensuring your approach to solving those problems work as cheaply as possible.
If you do a good job at all three phases, you win the luxury of doing it all again; this time with more complexity and scope. If you dont do well, you end up forfeiting or being asked to leave [the game](https://www.amazon.com/dp/B004W3FM4A/ref=dp-kindle-redirect?_encoding=UTF8&btkr=1).
## Problem discovery
The first phase of a planning cycle is exploring the different problems you could pick to solve. Its surprisingly common to skip this phase, but that unsurprisingly leads to inertia-driven local optimization. Taking the time to evaluate which problem to solve is one of the best predictors Ive found of a teams long-term performance.
The themes Ive found useful for populating the problem space are:
- **Users pain**. What are the problems that your users experience? Its useful to both go broad via survey mechanisms as well as to go deep by interviewing a smaller set of interesting folks across different user segments.
- **Users purpose**. What motivates your users to engage with your systems? How can you better enable them to accomplish their goals?
- **Benchmark**. Look at how your company compares to competitors in the same and similar industries. Are there areas that you are quite weak? Those are areas to _consider_ investment. Sometimes folks keep to a narrow lense when benchmarking, but Ive found that you learn the most interesting things by considering both fairly similar and rather different companies.
- **Cohorts**. What is hiding behind your clean distributions? Exploring your data for the cohorts hidden behind top-level analysis is an effective way to discover new kinds of users with surprising needs.
- **Competitive advantages.** By understanding the areas youre exceptionally strong in, you can identify opportunities that youre better positioned to fulfill than other companies.
- **Competitive moats**. Moats are a more extreme version of a competitive advantage. Moats represent a sustaining competitive advantage, which make it possible for you to pursue offerings that others simply cannot. Its useful to consider moats in three different ways:
- What your existing moats enable you to do today?
- What are the potential moats you could build for the future?
- What moats are your competitors luxuriating behind?
- **Compounding leverage**. What are the composable blocks that you could start building today that will compound into major product or [technical leverage](https://lethain.com/building-technical-leverage/) over time? I think of this category of work as finding ways to get the benefit (at least) twice. This are potentially tasks that initially dont seem important enough to prioritize, but whose compounding value makes it possible.
- A design example might be introducing a application new navigation scheme that better supports the expanded set of actions and modes you have today, and that will support future proliferation as well. (Bonus points if it manages to prevent future arguments about positioning of new actions relative to existing ones!)
- An infrastructure example might be moving a failing piece of technology to a new standard, this addresses a reliability issue, reduces maintenance costs, and also [reduces the costs of future migrations](https://lethain.com/migrations/).
## Problem selection
Once youve identified enough possible problems, the next challenge is to narrow down to a specific problem portfolio. Some of the aspects Ive found useful to consider during this phase are:
- **Surviving the round**. Thinking back to the iterative elimination tournament, what do you need to do to survive the current round? This might be the revenue the product will need to generate to avoid getting canceled, adoption, etc.
- **Surviving the next round**. Where do you need to be when the next round starts, to avoid getting eliminated then? There are a number of ways, many of them revolving around quality tradeoffs, to reduce long-term throughput in favor of short term velocity. (Conversely, winning leads to significantly more resources later, so that tradeoff is appropriate sometimes!)
- **Winning rounds**. Its important to survive every round, but its also important to eventually win a round! What work would ensure youre trending towards winning a round?
- **Consider different time frames.** When folks disagree which problems to work on, I find its most frequently rooted in different assumptions about the correct time frame to optimize for. What would you do if your company was going to run out of money in six months? What if there were no external factors forcing you to show results until two years out? Five years out?
- **Industry trends**. Where do you think the industry is moving towards, and what work will position you to take advantage of those friends, or at least avoid having to redo the work in near future?
- **Return on investment**. Personally, I think folks often under prioritize quick, easy wins. If youre in the uncommon position of understanding both the impact and costs of doing small projects, then take time to try ordering problems by expected return on investment. At this phase youre unlikely to know the exact solution, so figuring out cost is tricky, but for categories of problems youve seen before you can probably make a solid guess (if you dont personally have relevant experience, ask around). Particular in cases where wins are compounding, these are be surprisingly valuable over the medium and long term.
- **Experiments to learn**. What could you learn now that would make problem selection in the future much easier?
## Solution validation
Once youve narrowed down the problem you want to solve, its easy to jump directly into execution, but that can make it easy to fall in love with a difficult approach. Instead, Ive found it well worth it to derisk your approach with an explicit solution validation phase.
The elements Ive found effective for solution validation are:
- **Write a customer letter.** Write the launch announcement that you would send after finishing the solution. Are you able to write something exciting, useful and real? Its much more useful to test it against your actual users rather than relying on your intuition.
- **Identify prior art**. How do peers across the industry approach this problem? The fact that others have solved a problem in a certain way doesnt mean its a great way, but it does at least mean its possible. A mild caveat that its better to rely on folks you have some connection to instead of conference talks and such; there is a surprisingly large amount of misinformation out there.
- **Find reference users**. Can you find users who are willing to be the first users for the solution? If you cant, you should be skeptical whether what youre building is worthwhile.
- **Prefer experimentation over analysis**. Its far more reliable to get good at cheap validation than it is to get great at consistently picking the right solution. Even if youre brilliant, you are almost always missing essential information when you begin designing. Analysis can often uncover missing information, but it depends on knowing where to look, whereas experimentation allows you to find problems you didnt anticipate.
- **Find the path more quickly traveled**. The most expensive way to validate a solution is to build it in its entirety. The upside of that approach is that youve lost no time if you picked a good solution, the downside is that youve sacrificed a huge amount of time if its not. Try to find the cheapest way to validate.
- **Justify switching costs**. What will the switching costs be for users who move to your solution? Even if folks want to use it, if the switching costs are too high then they simply wont be able to. Test with your potential users if theyd be willing to pay the full cost of migrating to your solution instead of their existing planned work.
As an aside, Ive found that most aspects of [running a successful technology migration](http://lethain.com/migrations) overlap with good solution validation! This is a very general skill that will repay the time you invest into learning it many times over.
Putting these three elements todayexploration, selection and validationwont make you an exceptional product manager overnight, but they will provide a solid starting place to develop those skills and perspective for the next time you find yourself donning the product manager hat.

View File

@@ -0,0 +1,71 @@
Clipped from: [https://lethain.com/product-management-infra-engineering/](https://lethain.com/product-management-infra-engineering/)
![Discover Prioritize 000 ](Exported%20image%2020240808113927-0.png)
Recently a bunch of teams I work with have turned the corner, having paid down technical debt to a long-term sustainable level. The future unfurls with possibility. We can do _anything_. Thats exciting! It can also be pretty disorienting. For me, this is the most inspiring moment of management, and one of the hardest.
When we were completely focused on system reliability or churning tasks, most teams pulled their roadmaps down to a month or two, and we got so focused that we disconnected from our internal users. With less time soaked by maintenance, weve scurried to understand our users needs and define an optimistic, future-facing roadmap to support them.
In short, weve bootstrapped product management.
Many of the infrastructure engineer teams Ive been a part of have struggled to make the transition from maintenance to innovation, and I wanted to write down some of the ideas that were exploring to ease this shift. Id also love to hear what has worked well for other folks!
## Foundation to Innovation
I believe teams tend to have two distinct modes of operation:
- a foundation mode where the vast majority of tasks are mandatory, driven by non-negotiable needs like compliance, security, reliability and “victim of success” challenges like scaling a very popular product. Kanban is optimized for this mode, and I think [The Phoenix Project](https://www.amazon.com/dp/B00AZRBLHO/ref=dp-kindle-redirect?_encoding=UTF8&btkr=1) is a really helpful resource on executing in this mode.
- an innovation mode where you have a lot of flexibility in which problems to prioritize and how to solve them. This is similar to product development as [described in Inspired](https://www.amazon.com/INSPIRED-Create-Tech-Products-Customers/dp/1119387507/ref=dp_ob_image_bk).
In practice many teams have one foot in both modes, and most teams cycle between the two over time, but for any given team at any given time, they usually have a primary model.
The foundation mode is a game of execution, focus and limiting work-in-progress, whereas the innovation mode is about listening to users, exploring solution spaces and an eternal focus on validating solutions as early and cheaply as possible.
Most infrastructure teams have a lot of experience in foundation, but have much less in innovation, so I wont belabor managing through scarcity, and will instead dive into how to manage in times of surplus engineering capacity.
## Problem discovery
When you have surplus engineering capacity, folks tend to have a long backlog of stuff theyd like to work on, and many teams immediately jump on those, but I think its useful to fight that instinct and to step back and do deliberate discovery.
There are two things that are essential to discovering opportunity: cast a very broad net, and dont evaluate ideas while discovering them. Try to get as many ideas as possible, and dont spend a single moment worrying if theyre any good, prioritization is a later step.
I recommend taking a four prong approach, learning about your users needs, what peer companies are doing, where industry leaders are going, and brainstorming with your team.
Some of the techniques that Ive seen work well:
- **SLAs** - sit down with your users to discuss what SLAs they expect from your system, and then turn those into your dashboards. Even if you cant hit some of those SLAs today, understanding what users want you to be able to offer is powerful.
- **User surveys** - its surprisingly hard to write a great survey, and easy to create survey fatigue, but a good survey is an awesome way to get many folks to share input. A good survey is a short, quantifies when possible, and gets proofread by folk with opposing perspectives before its sent out! Theyre especially good as a first step to identify who to follow up with in detail.
- **Coffee chats** - meet with your users periodically and learn about what theyre doing. Do remember to ask if there is stuff you can improve, but I think its most valuable to understand what theyre doing, which is great fodder for thinking about how you can help. A cheaper version of this is a short email with a quick compliment and asking if you could help with anything.
- **Discussion groups** - bringing together a small group of two to four folks and hearing their ideas is a good way to get input. I find these work best when you can bring a specific proposal, or set of proposals, for folks to react against. As a variant, weve also experimented with recurring customer advisory groups.
- **Peer-company chats** - in addition to chatting with your users, Ive found it equally valuable to chat with a wide variety of folks working on similar problems at other companies to understand how theyre thinking about things. Ive heard some extremely valid concerns around this leads directly to cargo-culting and groupthink, which is why I think its so important to decouple discovery from prioritization: listen to what other folks are doing, dont automatically decide to adopt them.
- **Academic and industry research** - look for papers coming out of the big tech companies (Google, Microsoft, Facebook, etc) and also for academic research on a given area. I did a basic example of this [when I surveyed load generation research](https://lethain.com/braindump-on-load-generation/), and it was helpful to expand my thinking beyond the obvious.
- **Open source** - [along the lines of my exploration of the open source data ecosystem](https://lethain.com/from-lambda-to-kappa-dataflow-paradigms/), I find it useful to spend some time understanding the state of open source for a given area. This gives you an easy sense of trends and how other folks see the space evolving.
- **Cloud vendors** - with [cloud offerings rapidly expanding](https://lethain.com/physics-of-cloud-expansion/), its useful to take a look at what new related offerings have popped up on AWS, Azure, GCP and such.
The goal of each of these is to collect a tremendous amount of information. My mental model of this phase is to load as much state into your head as possible, to give you a broad perspective as you move into prioritization.
## Prioritization
Brimming with research and user needs, I try to create three artifacts:
- A **charter** that identifies the unique value your team tries to provide, your competitive advantages and the strategy that youll employ.
- An **optimistic three-year vision** of the best possible system youd like to be providing. Constrain the vision by what is possible, but dont constraint it by what is reasonable. My rule of thumb is “where could we be if everything went perfectly for three years?”
- A **prioritized list of user pain** that captures your users active needs, and in particular buckets the concerns together into things that you might be able to address, eliminate or empower with a single solution.
Your aim is to relieve as much user pain as possible while also making forward progress towards your optimistic vision. I think doing this well is the artistry of product management, and I believe each project can usually make significant progress towards both.
Changing priorities late in a project is expensive, so we try to align on priorities as early as possible, and allow users as much input as possible early to weigh in on which parts will be useful to them, and in particular which parts can be used independently. A great list of priorities allows us to deliver incremental value to our users at each phase, not waiting until the last step to deliver a big bang of utility.
Sometimes current needs dont align well with the future, or the future is a long ways away, and in that case my rule of thumb is to invest 70% of your effort on solving immediate user needs, and 30% on advancing towards the future. I advocate 70% towards immediate user needs, because I think folks tend to over index on the future, and this helps avoid falling into that trap. [In areas where cloud vendors are rapidly expanding](https://lethain.com/physics-of-cloud-expansion/), sometimes it may be reasonable to devote 100% of your efforts to immediate user pain, on the assumption that cloud offerings will be available in the next 2-3 years to absorb the operational load and technical debt of the existing solution. Some areas are not likely to get cloud investment in the near term, but increasingly we should be looking at areas we can strategically underinvest today on the assumption that the clouds will provide an easy solution soon.
## Solution validation
Once youve decided what to focus on solving, a many teams immediately go “full waterfall”, designing a twelve month roadmap. Long-term planning is hard to resist, because pretty much every user and every planning process demands it of you, so youll probably end up writing an artifact of this nature, but I beg you not to pretend it means something.
This isnt because estimating is hardalthough estimating is hardbut because it assumes that our solutions are actually good solutions. Its my opinion that most solutions are, in fact, pretty bad. The secret is not to “get brilliant” and pick better solutions, but rather to get skeptical and to design your approach to validate the value and practicality of solutions as cheaply as possible.
I often assume that internal users have a great sense of existing capabilities, but that sometimes isnt the case. Part of “getting skeptical” is evangelizing your existing solutions and seeing if there is already a reasonable way to solve the problem under discussion that maybe you havent documented well or requires some architectural familiarity between both the users and the providers to realize a solution is already available.
Its a common pattern to do the hardest migration first, and Im a big fan of that approach, but even then youve spent the majority of the time required to build a system by the time youre validating it. Do experiments, gather data to prove the approach wont work, validate the approach has worked for others. Try as hard as possible to prove your solution cannot work.
Once youve validated a solution, you also want to keep in mind similar needs from other users. The pattern might look like doing one hard integration first, and then doing a staccato burst of easy integrations to smooth the edges and check for broad applicability.
Weve been experimenting with the pattern of embedding the team building a solution into the team theyre building the solution for, and that feels like a good strategy for quick iteration and avoiding falling in love with awesome things that are not necessarily useful things.
## Closing
There are many infrastructure engineering teams which dont make the transition from maintenance to innovation, and some which intentionally decide against doing so. Its an uncomfortable transition, but Ive found it remarkably rewarding: more direct contribution to your coworkers success, more excitement from other leaders within the company, and shrugging off the mantle of cost center to become an acknowledged source of innovation.
While Im really excited at how this approach has helped us focus on directly supporting our users, its still an early approach. Three questions in particular continue to bounce around in my head:
1. User discovery takes up a bunch of time from other teams. How can we be more effective with their time?
2. How do we avoid falling in love with solutions, particularly for those that are difficult to validate early?
3. How do other folks do this!?
If youre doing something different, or even if youre doing something similar, Id love to hear from you!
Thanks to [Amy](https://twitter.com/amyngyn), [Peter](https://twitter.com/peterseibel) and [Ranbir](https://twitter.com/TheRanbirChawla) for shaping this post.
Published on February 6, 2018.

View File

@@ -0,0 +1,91 @@
Clipped from: [https://fennel.ai/blog/real-world-recommendation-system/](https://fennel.ai/blog/real-world-recommendation-system/)
Training a collaborative filtering based recommendation system on a toy dataset is a sophomore-year project in colleges these days. But where the rubber meets the road is building such a system at scale, deploying in production, and serving live requests within a few hundred milliseconds while the user is waiting for the page to load. To build a system like this, engineers have to make decisions spanning multiple moving layers like:
- High-level paradigms (like collaborative filtering, content based recommendations, vector search, model based recommendations)
- ML algorithms (e.g., GBDTs, SVD, Multi tower neural networks, etc.)
- Modeling libraries (e.g., PyTorch, Tensorflow, XGBoost)
- Data management (e.g., choice of DB, caching strategy, reuse primary database or copy all the data in another system optimized for recommendation workload, etc.) [[1]](https://fennel.ai/blog/real-world-recommendation-system/#fn1)
- Feature management (e.g., offline vs online, precompute vs serve live)
- Serving systems (performance, query latency, distribution model, fault tolerance, etc.)
- Deployment system (e.g., how does new code get updated, build steps, keeping caches working after processes restart, etc.)
- Hardware (e.g., GPUs, SSDs)
No wonder architecting a system like this is a daunting task. Thankfully though, after years of trial and error, FAANG and other top tech companies have independently converged on a common architecture for building/deploying production-grade recommendation systems. Further, this architecture is domain/vertical agnostic and can power all sorts of applications under the sun — from e-commerce and feeds to search, notifications, email marketing, etc.
The goal of this publication is to start from the basics, explain nuances of all the moving layers, and describe this universal recommendation system architecture.
We will start with a post to explain the serving side of this architecture at a high level, quickly followed by a post about the training side — these two posts will mostly outline the structure and identify the key scaling problems in serving and training, respectively. (Edit, we ended up writing two follow-ups to this post — part 2 about [training data generation here](https://fennel.ai/blog/real-world-recommendation-systems/) and [part 3 about modeling here](https://fennel.ai/blog/real-world-recommendation-systems-21e/)). Future posts will go through these scaling issues one by one and describe how they are typically solved, along with the best practices developed over years of learning. So lets get started:
Modern recommendation systems are composed of eight (somewhat overlapping) logical stages:
1. Retrieval
2. Filtering
3. Feature Extraction
4. Scoring
5. Ranking
6. Feature Logging
7. Training Data Generation
8. Model Training
The first five of these are related to serving, and the last three are related to training.
![https://www.fennel.ai/blog/content/images/2022/10/1cacae00-cc59-46b1-bac2-71eeb72828ac_960x720-1.jpeg](Exported%20image%2020240808113928-0.jpeg)
Lets go through all the serving layers one by one:
## **1. Retrieval**
Products like Facebook have millions of things to show in any recommendation unit - so many that its physically impossible to score all of them using any ML model while the user is waiting for their feed to load. So instead of scoring each item in the inventory, a more manageable subset of the inventory is first obtained via a process called Retrieval or “Candidate Generation” (since it generates candidates for ranking).
Retrieval is not just a FAANG scale problem — since the user is waiting for the “page” to load, most recommendation requests have a budget of only 500ms or so, and it is only possible to score a few hundred items in a request. As a result, whenever the inventory is a couple thousand items or more (which covers a large % of all real-world systems), a retrieval phase is needed.
How does Retrieval work? Retrieval is done by writing a few heuristics, also called as “candidate generators” or simply “generators”, each of which selects, say a dozen or so distinct candidates. Some common examples of generators are:
- Content that is trending in a users geography in the last x hours
- Recent content from authors/topics that the user explicitly “follows”
- Find 5 contents that user “liked” in the past, and for each such content, find 5 more “related” items
- Find the most relevant topics for a user and find the freshest content from each of the topics.
Retrieval can be powered by ML (e.g., trained embeddings), but more often than not, a larger % of generators are mere heuristics that encode some “product thinking” about what content is likely to create a good recommendation experience. And by writing a few of these and taking a union of all their candidates, we ensure that the system is able to at least consider all sorts of interesting inventory. Retrieval has only two jobs — 1) get all the interesting things (or at least as many as possible) [[2]](https://fennel.ai/blog/real-world-recommendation-system/#fn2) and 2) get as few total things as possible so that we can score/examine each candidate using the power of ML.
## 2. Filtering
After retrieving a few hundred candidates, recommendation systems typically filter out “invalid inventory”. For instance, if youre building a social network, you might want to filter out things that are likely to be spammy. Or if you are building a video OTT platform, you may have to do some geo-licensing-based filtering. Or if youre building an e-commerce product, you may have to filter things that are out of stock. Filters can also be extremely personalized, some examples:
1. Some products try to filter out content that the user has already seen before
2. Some products expose some controls to the users to hide away topics or authors or other sources of content
In short, most real-world recommendation systems develop a long list of filters over time, which once again encode some product thinking about what creates a good experience. [[[3]](https://fennel.ai/blog/real-world-recommendation-system/#fn3)
Filtering and retrieval have a very interesting relationship. Some filters are pushed down to the generators themselves — for instance, if youre building a dating product, filters for location and sexual preferences may be a part of each generator itself. But more often than not, it is physically impossible to have each generator respect each filter at the source, and so a whole layer of filtering is needed.
## 3. Feature Extraction
After filtering, we have a slightly smaller list of candidates — but its still going to be a couple hundred candidates long. We somehow need to choose the top ten items to show to the user. As you can imagine, this is going to involve some sort of scoring — for instance, in a job portal, we may want to compute how close is the jobs salary range to the users desired salary range.
But before any scoring can even begin, we need to obtain a bunch of data about each item that is going to be scored. In the job portal example, we will need the salary range of each candidate's job. Such signals about candidates are called “features”. Features are not just about the candidates but also include data about the user (e.g., users desired salary range). In fact, some of the most important features in literally every single recommendation system are those that capture users interaction behavior with potential candidates - this is so important and so nuanced that we will dedicate a whole post on this topic in this blog soon. Either way, we have a few hundred candidates, and we extract a bunch of features (which are basically just pieces of data) about the candidates and the user. Usually, we get anywhere between a few dozen to hundreds of features per candidate.
It is worth pausing here and letting the scale sink in for a minute - we have a few hundred candidates, say 1000, and we get a few dozen features about each candidate, say 100 — we need to fetch 1000x100 or 100K pieces of data from some database in only 500ms (the latency budget of a recommendation request). And note that you dont even have to be at FAANG scale to run into this problem - even if you have a small inventory (say a few thousand items) and a few dozen features, youd still run into this problem. And fetching and computing so much data is an incredibly hard infra problem to solve and as a result, creates huge limitations on how “expressive” the features can actually be.
## 4. Scoring
So far, we have narrowed down the full inventory to a few hundred candidates and extracted a few dozen features about each. Now comes the bit where we use all the extracted features to assign a score to each candidate. In the simplest systems, the scoring phase is pretty rudimentary, often just a handcrafted formula that mixes a bunch of features of interest (e.g “lets divide the number of likes by the number of impressions and give a boost by doubling the score if the user follows the content author”). But very soon, these handcrafted formulae and rules start hitting “corner cases” and creating bad experiences. That is where ML kicks in - a machine learning model is trained that takes in all these dozens of features and spits out a score (details on how such a model is trained to be covered in the next post).
There are two key ideas that are very successful and present in the scoring of most real-world recommendation systems (and wed write dedicated posts about both in the future — stay tuned):
1. Multi-stage scoring — not all ML models are equal, and some are lot “heavier” than others. And it is usually not possible to run the heaviest ML models on hundreds of candidates. So instead, scoring itself is broken down in two substages — 1st stage scoring (which uses a relatively lighter ML model like GBDTs on all 500 candidates and emits out, say, top 100 candidates) and the 2nd stage scoring, which runs the heavy model (say deep neural network) on just the top 100 candidates.
2. Combining many models — ML models can only learn whatever we teach them to learn. And typically, they are taught to predict the probability of user engaging in a single action, say like. Sorting all content by what gets clicked is a good start but has lots of issues — for instance, it might only distribute clickbaity content. To make the recommendations more balanced, usually, multiple models are trained - say, one for predicting clicks, one for predicting comments, one for user reporting the content, etc. And the final score of a candidate is a weighted average of all these models. While this makes the recommendations better, this can also increase the amount of computation that needs to be done. [[4]](https://fennel.ai/blog/real-world-recommendation-system/#fn4)
## 5. Ranking
Once scores have been computed for every candidate, the system moves on to the very last step — ranking. In the simplest systems, this stage is as simple as sorting all the candidates on their scores and just taking the top K. But in more complicated systems, the scores themselves are perturbed using non-ML business rules. For instance, it is a common requirement across many products to diversify the results a bit — for instance, not show content from the same publisher/author one after another. There are many algorithms for such diversification, but most of them operate in a similar fashion by adjusting the scores to respect the diversity (e.g., demote scores if successive items are not diverse enough).
In addition to score perturbation, it is a common practice to run all the items against all the filters once again at this stage to avoid any embarrassing failures. For instance, maybe some of the candidate generators are a bit stale and dont know that an item has gone out of stock — its better to filter it out here instead of sending an item to the user that they cant even purchase.
Finally, once top K items are chosen, they are handed to some sort of “delivery” system which is responsible for things like pagination, caching, etc.
## Conclusion
Thats it! These are the five serving stages of real-world recommendation systems. As you can see, even if a recommendation system is trained, deploying that in production is incredibly hard. In the next post, we will look at the training side of the recommendation system. And in the subsequent posts, we will go through all the infra/scaling issues outlined here and share how they are typically solved — stay tuned!
1. Managing data is surprisingly hard for real-world recommendation systems because of extreme needs on all three of write throughput, read throughput, and read latencies. As a result, primary databases (e.g., MySQL, MongoDB, etc.) almost never work out of the box (unless you put in a lot of work to scale them in a clever way) [↩︎](https://fennel.ai/blog/real-world-recommendation-system/#fnref1)
2. Often bloom filters are used to filter out content already “seen” by the users [↩︎](https://fennel.ai/blog/real-world-recommendation-system/#fnref2)
3. This is actually fairly true across all the stages — 80% of decisions & iterations in building a recommendation system are all about product-specific needs & business rules, not machine learning. [↩︎](https://fennel.ai/blog/real-world-recommendation-system/#fnref3)
4. This technique is called “value modeling” - value model is an expression (typically weighted linear sum) of ML models with each term describing a specific kind of value to the user. [↩︎](https://fennel.ai/blog/real-world-recommendation-system/#fnref4)
### We publish high-quality technical content on all things machine learning.
_Subscribe to our blog to get updates delivered right to your inbox._
Please wait...
Please check your inbox and click the link
Please enter a valid email address!

View File

@@ -0,0 +1,14 @@
Clipped from: [https://github.com/StackExchange/StackExchange.Redis/blob/main/docs/ThreadTheft.md](https://github.com/StackExchange/StackExchange.Redis/blob/main/docs/ThreadTheft.md)
If you're here because you followed a link in an exception and you just want your code to work, the short version is: try adding the following _early on_ in your application startup:
ConnectionMultiplexer.SetFeatureFlag("preventthreadtheft", true);
and see if that fixes things. If you want more context as to what this is about - keep reading!
Behind the scenes, for each connection to redis, StackExchange.Redis keeps a queue of the commands that we've sent to redis that are awaiting a reply. As each reply comes in we look at the next pending command (order is preserved, which keeps things simple), and we trigger the "here's your result" API for that reply. For async/await code, this then leads to your "continuation" becoming reactivated, which is how your code comes back to life when an await-ed task gets completed. That's the simple version, but reality is a bit more nuanced.
By _default_, when you trigger TrySetResult (etc) on a Task, the continuations are invoked _synchronously_, i.e. the thread that is setting the result now goes on immediately to run whatever it is that your continuation wanted. In our case, that would be very bad as that would mean that the dedicated reader loop (that is meant to be processing results from redis) is now running your application logic instead; this is **thread theft**, and would exhibit as lots of timeouts with rs: CompletePendingMessage in the information (rs is the **r**eader **s**tate; you shouldn't often observe it in the CompletePendingMessage* step, because it is meant to be very fast; if you are seeing it often it probably means that the reader is being hijacked when trying to set results).
To _avoid_ this, we use the TaskCreationOptions.RunContinuationsAsynchronously flag. What _this_ does depends a little on whether you have a SynchronizationContext. If you _don't_ (common for console applications, services, etc), then the TPL uses the standard thread-pool mechanisms to schedule the continuation. If you _do_ have a SynchronizationContext (common in UI applications and web-servers), then its Post method is used instead; the Post method is _meant_ to be an asynchronous dispatch API. But... not all implementations are equal. Some SynchronizationContext implementations treat Post as a synchronous invoke. This is true in particular of LegacyAspNetSynchronizationContext, which is what you get if you configure ASP.NET with:
<add key="aspnet:UseTaskFriendlySynchronizationContext" value="false" />
or if you do _not_ have a <httpRuntime targetFramework="..." /> of at least 4.5 (which causes the above to default true) like this:
<httpRuntime targetFramework="4.5" />
([citation](https://devblogs.microsoft.com/aspnet/all-about-httpruntime-targetframework))
In these scenarios, we would once again end up with the reader being stolen and used for processing your application logic. This can doom any further awaits to timeouts, either temporarily (until the application logic chooses to release the thread), or permanently (essentially deadlocking yourself).
To avoid this, the library includes an additional layer of mistrust; specifically, if the preventthreadtheft feature flag is enabled, we will _pre-emptively_ queue the completions on the thread-pool. This is a little less efficient in the _default_ case, but _if and only if_ you have a misbehaving SynchronizationContext, this is both appropriate and necessary, and does not represent additional overhead.
The library will attempt to detect LegacyAspNetSynchronizationContext in particular, but this is not always reliable. The flag is also available for manual use with other similar scenarios.

View File

@@ -0,0 +1,66 @@
Clipped from: [https://www.productplan.com/glossary/scrumban/](https://www.productplan.com/glossary/scrumban/)
## What Is Scrumban?
Scrumban is a project management framework that combines important features of two popular agile methodologies: Scrum and Kanban. The Scrumban framework merges the structure and predictable routines of Scrum with Kanbans flexibility to make teams more agile, efficient, and productive.
For companies that implement Scrumban, the approach can help their teams focus on the correct strategic tasks while at the same time improving their processes.
## How Does Scrumban Combine Scrum and Kanban?
To understand how Scrumban merges Scrum and Kanban, we first need to understand each of these frameworks.
## The Basics of Scrum
![The Basics of Scrum](Exported%20image%2020240808113924-0.png)
A scrum is an [agile](https://www.productplan.com/agile-product-management/) approach used in software development. With Scrum, a team organizes itself into specific roles, including a Scrum master, product owner, and the rest of the Scrum team. The team breaks its workload into short timeframes called [sprints](https://www.productplan.com/glossary/sprint/). Each sprint lasts two weeks or one month.
During a sprint, the developers work only on the tasks the team agreed to during the [sprint meeting](https://www.productplan.com/glossary/sprint-planning/). Before the next sprint, the team holds another sprint meeting and decides which items to work on next. Scrum teams also meet each morning for short [standups](https://www.productplan.com/glossary/standup/) to discuss the days tasks.
## The Basics of Kanban
Kanban is a visual approach to managing a teams workload. With this methodology, a team creates a Kanban board to visually display its workflow in columns—such as “Ready to Start,” “In Progress,” “Under Review,” and Completed.”
As developers begin working on an item, they move a card (or sticky note) with the items name from the Ready-to-Start column to In-Progress. If an item needs to move backward, from Under Review back to In-Progress—the team can move that card back to the In-Progress column. The Kanban board makes it easy for everyone to view and update the status of each project quickly.
## The Basics of Scrumban
Scrumban merges the structure and predictability of Scrum with Kanbans flexibility and continuous workflow. When implemented correctly, Scrumban can help a team benefit from both the prescriptive nature of Scrum and the freedom of Kanban to improve their processes.
## How Does Scrumban Work? (A Step-by-Step Guide)
![How Does Scrumban Work | ProductPlan](Exported%20image%2020240808113924-1.png)
Scrumban involves applying Kanban principles—visualization of workflow, and flexible processes—to a teams Scrum framework. But, Scrumban removed some of the more rigid aspects of Scrum and left each team to create a custom approach to development.
Here is a step-by-step guide to developing a Scrumban framework for your team.
### Step 1: Develop a Scrumban board
A Scrumban board is similar to a Kanban board. Because you will be using it as your primary workflow tool, add as many columns to your Scrumban board as your team needs to mark each discrete phase of progress. But be careful not to create so many columns that the board becomes cluttered and difficult to view.
### Step 2: Set your work-in-progress limits
Remember, Scrum sets both time and task limits for every sprint. Kanban, by contrast, focuses on continuous workflow. You will need to establish a limit on how much work your team can take on at any one point. For Scrumban, that limit will be the number of total cards on the board at any time. Set a realistic limit to keep your team from becoming overwhelmed and frustrated.
### Step 3: Order the teams priorities on the board
This step illustrates another major difference between Scrum and Kanban (and Scrumban). With Scrum, youll assign tasks to specific individuals within your dev group for each sprint. Under Scrumban, on the other hand, your focus will be establishing the priority order of all projects on the board. Your team will decide which person will tackle which tasks.
### Step 4: Throw out your planning-poker cards
Because each sprint has a strict time limit, and the team can work on only a pre-defined set of projects during any one sprint, a Scrum team needs to estimate how long each development task will take. So theyve devised methods such as [planning poker](https://www.productplan.com/glossary/planning-poker/) to estimate the number of [story points](https://www.productplan.com/glossary/story-point/) (indicating time and difficulty) for each task. With Scrumban, work is continuous and not time-limited, so your team wont estimate story points. Youll focus only on [prioritizing](https://www.productplan.com/product-management-frameworks/) the most important projects.
### Step 5: Set your daily meetings
Although you wont have most of the meetings typical of the Scrum framework—sprint planning, sprint review, retrospective—Scrumban meetings can include short standups for the team to discuss their plans and challenges for the upcoming day. These short meetings are also a good way to encourage team bonding and cohesion because your developers will spend a lot of time working individually on their tasks and might not have much time for interaction otherwise.
### [Get Strategic Project Alignment ➜](https://go.productplan.com/cs/c/?cta_guid=5e81c0e5-ef7c-4934-9a09-1ffdb75ef256&signature=AAH58kEKHQbQJW6629FIXlrtbkdl61qPTg&placement_guid=bfb5032e-5746-4c05-9f2a-54b36ba0e871&click=50d2ee6d-3f9f-4e1f-9706-42febdab50e8&hsutk=&canon=https%3A%2F%2Fwww.productplan.com%2Fglossary%2Fscrumban%2F&portal_id=3434168&redirect_url=APefjpH62_akcPKt5r-wAyIGTPz-MzD-4jlmpUTA08t06kdMM_hG7YDHiJQk--6f3mPj8oHX8vTPELynngVJ9O9d_7KIARUO-JocvlEzt2M9qjELNmVc2lpJJElNxTAKuqZh8mlczCZJFlgGsHThwFBq0XDen52xBw)
## When Should a Team Use Scrumban?
A team can benefit from the Scrumban approach under several circumstances. For example:
**1.** **For maintenance of ongoing projects.**
These could include projects in which, unlike a new product launch, there is no definitive completion date for the work.
**2. For a team having trouble with Scrum.**
It can happen for several reasons. For example, the company doesnt have enough resources to support a Scrum environment, or the team finds Scrums requirements too rigid.
**3. When a company wants to give its team more flexibility in how it works.**
With Scrum, the team often assigns specific tasks to individuals for each sprint. But Scrumban only sets a broad list of projects and lets the team itself determine how best to leverage its resources. It enhances teamwork and enables individuals in the company to find the projects best suited to their skills and interests.
**Related Terms:** [Scrum agile framework](https://www.productplan.com/glossary/scrum-agile-framework/) / [Kanban board](https://www.productplan.com/glossary/kanban-board/) / [Kanban roadmap](https://www.productplan.com/glossary/kanban-roadmap/) / [Scrum master](https://www.productplan.com/glossary/scrum-master/) / [LeSS (Large Scale Scrum)](https://www.productplan.com/glossary/less-large-scale-scrum/)

View File

@@ -0,0 +1 @@
![[Attachment.pdf]]