Pete Cheslock
@petecheslock.com
๐ค 816
๐ฅ 61
๐ 26
๐ฅฉ He/Him ๐ "Anything worth doing is worth overdoing."
reposted by
Pete Cheslock
Yuan Tang
13 days ago
Most teams over-engineer their inference stack from day one. They disaggregate before measuring. Add speculative decoding before concurrency stabilizes.
5
9
1
reposted by
Pete Cheslock
Yuan Tang
20 days ago
Part 2 of our ๐๐ถ๐๐๐ฟ๐ถ๐ฏ๐๐๐ฒ๐ฑ ๐๐ ๐๐ป๐ณ๐ฒ๐ฟ๐ฒ๐ป๐ฐ๐ฒ series is now live on Red Hat Developer: ๐๐ฑ๐ต๐ช๐ฎ๐ช๐ป๐ช๐ฏ๐จ ๐๐ช๐ด๐ต๐ณ๐ช๐ฃ๐ถ๐ต๐ฆ๐ฅ ๐๐ ๐๐ฏ๐ง๐ฆ๐ณ๐ฆ๐ฏ๐ค๐ฆ: ๐๐ฅ๐ท๐ข๐ฏ๐ค๐ฆ๐ฅ ๐๐ฆ๐ฑ๐ญ๐ฐ๐บ๐ฎ๐ฆ๐ฏ๐ต ๐๐ข๐ต๐ต๐ฆ๐ณ๐ฏ๐ด. In Part 1, we covered prefill/decode phases and the 5D parallelism framework.
4
9
2
reposted by
Pete Cheslock
Yuan Tang
24 days ago
Excited to share Part 1 of our blog series on Red Hat Developer: ๐๐ฆ๐ด๐ช๐จ๐ฏ๐ช๐ฏ๐จ ๐๐ช๐ด๐ต๐ณ๐ช๐ฃ๐ถ๐ต๐ฆ๐ฅ ๐๐ ๐๐ฏ๐ง๐ฆ๐ณ๐ฆ๐ฏ๐ค๐ฆ: ๐๐ฐ๐ณ๐ฆ ๐๐ฐ๐ฏ๐ค๐ฆ๐ฑ๐ต๐ด ๐ข๐ฏ๐ฅ ๐๐ค๐ข๐ญ๐ช๐ฏ๐จ ๐๐ช๐ฎ๐ฆ๐ฏ๐ด๐ช๐ฐ๐ฏ๐ด. LLM inference is two workloads pretending to be one. The prefill phase is compute-bound, processing entire prompts in parallel to populate the KV cache.
2
6
2
That time I gave some quick legal advice to Afroman before he won his court case.
4 months ago
0
5
0
Hey Boston friends, we're cooking up another great event in the area. Workshop + evening sessions covering: - vLLM project update - Model compression and speculative decoding - Agentic AI with vLLM - Distributed inference at scale with llm-d and k8s
luma.com/4rmkrrb7
loading . . .
vLLM Inference Meetup ยท Boston ยท Luma
Deep technical sessions. Live demos. Real conversations. If you're deploying, or scaling LLM inference, this is the room to be in. Join Red Hat AI, IBM,โฆ
https://luma.com/4rmkrrb7
4 months ago
0
1
1
reposted by
Pete Cheslock
Yuan Tang
4 months ago
๐ข ๐ง๐ต๐ฒ ๐ฆ๐๐ฎ๐๐ฒ ๐ผ๐ณ ๐ ๐ผ๐ฑ๐ฒ๐น ๐ฆ๐ฒ๐ฟ๐๐ถ๐ป๐ด ๐๐ผ๐บ๐บ๐๐ป๐ถ๐๐ถ๐ฒ๐: ๐ ๐ฎ๐ฟ๐ฐ๐ต ๐๐ฑ๐ถ๐๐ถ๐ผ๐ป ๐ถ๐ ๐ผ๐๐! We launched our newsletter publicly last year to share our contributions to upstream communities from our Red Hat AI teams. Weโve gained over ๐ญ๐ฏ๐ฌ๐ฌ ๐๐๐ฏ๐๐ฐ๐ฟ๐ถ๐ฏ๐ฒ๐ฟ๐!
2
2
2
reposted by
Pete Cheslock
Fred Katz
4 months ago
Jayson Tatum looking like Jayson Tatum is a horrifying development for the rest of the East.
1
53
6
I'm going to be in NYC next week, come and join me at the first llm-d meetup. If you're looking to learn more about distributed inferencing on kubernetes, this is going to be the place to be.
add a skeleton here at some point
5 months ago
0
0
0
reposted by
Pete Cheslock
llm-d
5 months ago
In the latest llm-d release, weโre tackling high hardware costs with the new GPU Recommendation Tool! ๐ Evaluate throughput, latency, and cost-effectiveness before requesting expensive cluster resources. Check out the full demo:
www.youtube.com/watch?v=Y26i...
loading . . .
Optimizing LLM Workloads: A Deep Dive into the GPU Recommendation Tool & Configuration Explorer
YouTube video by llm-d Project
https://www.youtube.com/watch?v=Y26i69zI6Ag
0
2
1
Come and join us for the first llm-d meetup in NYC!
add a skeleton here at some point
5 months ago
0
1
0
reposted by
Pete Cheslock
llm-d
5 months ago
The agenda is still evolving, and weโve got even more awesomeness in the works! ๐ Whether you're running GenAI in production or building the platforms to support it, this is the room to be in. ๐ March 11 | 4:30 PM ๐ 1 Madison Ave, NYC ๐๏ธ RSVP:
luma.com/0crwqwg4
loading . . .
Distributed Inference Meetup NYC ยท Luma
llm-d Distributed Inference Meetup NYC Hosted by Red Hat AI, IBM Research, and AMD, this event takes place on March 11, 2026 in New York City. What toโฆ
https://luma.com/0crwqwg4
0
0
2
reposted by
Pete Cheslock
Yuan Tang
5 months ago
We'd like to announce that
@kubernetes.io
WG Serving has succeeded and will be disbanded! Thank you everyone who have participated and contributed to the discussions and initiatives! More details:
groups.google.com/a/kubernetes...
loading . . .
[Announcement] WG Serving Has Succeeded and Will Be Disbanded
https://groups.google.com/a/kubernetes.io/g/dev/c/nDjMph1146A
1
4
3
reposted by
Pete Cheslock
llm-d
5 months ago
In case you missed it, last week the llm-d community shipped the v0.5 release. Check out the post from the llm-d project owners to learn more about all the features we've included in this release.
llm-d.ai/blog/llm-d-v...
add a skeleton here at some point
0
1
1
reposted by
Pete Cheslock
Yuan Tang
5 months ago
๐ข ๐ง๐ต๐ฒ ๐ฆ๐๐ฎ๐๐ฒ ๐ผ๐ณ ๐ ๐ผ๐ฑ๐ฒ๐น ๐ฆ๐ฒ๐ฟ๐๐ถ๐ป๐ด ๐๐ผ๐บ๐บ๐๐ป๐ถ๐๐ถ๐ฒ๐: ๐๐ฒ๐ฏ๐ฟ๐๐ฎ๐ฟ๐ ๐๐ฑ๐ถ๐๐ถ๐ผ๐ป ๐ถ๐ ๐ผ๐๐! We launched our newsletter publicly last year to share our contributions to upstream communities from our Red Hat AI teams. Weโve gained over ๐ญ๐ฎ๐ฌ๐ฌ ๐๐๐ฏ๐๐ฐ๐ฟ๐ถ๐ฏ๐ฒ๐ฟ๐!
1
1
1
reposted by
Pete Cheslock
llm-d
5 months ago
๐๏ธ llm-d v0.5: Sustaining Performance at Scale In our last release, we focused on breaking latency records. With v0.5, weโre shifting from peak performance to the operational rigor required to sustain those gains in production. ๐งต๐
llm-d.ai/blog/llm-d-v...
loading . . .
llm-d 0.5: Sustaining Performance at Scale | llm-d
Announcing the llm-d 0.5 release
https://llm-d.ai/blog/llm-d-v0.5-sustaining-performance-at-scale
1
1
2
reposted by
Pete Cheslock
llm-d
6 months ago
Standardizing high-performance inference requires deep ecosystem collaboration. ๐ Huge shoutout to @vllm_project and @IBMResearch on the new KV Offloading Connector. Weโre seeing up to 9x throughput gains on H100s and massive TTFT reductions. ๐งต
blog.vllm.ai/2026/01/08/k...
loading . . .
Inside vLLMโs New KV Offloading Connector: Smarter Memory Transfer for Maximizing Inference Throughput
In this post, we will describe the new KV cache offloading feature that was introduced in vLLM 0.11.0. We will focus on offloading to CPU memory (DRAM) and its benefits to improving overall inferenceโฆ
https://blog.vllm.ai/2026/01/08/kv-offloading-connector.html
1
0
1
reposted by
Pete Cheslock
llm-d
6 months ago
AI inference is like a busy airport: without a controller, you get gridlock. โ๏ธ Check out this breakdown by Cedric Clyburn from Red Hat on how llm-d intelligently routes distributed LLM requests. ๐น Solves "round robin" congestion ๐น Disaggregates P/D to save costs
www.youtube.com/watch?v=CNKG...
loading . . .
LLMโD Explained: Building NextโGen AI with LLMs, RAG & Kubernetes
YouTube video by IBM Technology
https://www.youtube.com/watch?v=CNKGgOphAPM
0
1
1
If you stop to think about it, Geysers are just Earth farts.
about 1 year ago
0
2
1
So a long time ago when buying new headphones and reading reviews, I noticed how the reviews often sounded similar to reviews for a bottle of wine. Like: "Rich and full-bodied with excellent depth. The bass notes are particularly impressive, with a smooth finish that lingers pleasantly."
over 1 year ago
2
2
1
reposted by
Pete Cheslock
worst guy you know
about 3 years ago
43
6657
1799
How do you pronounce โwwwโ the abbreviation for โWorld Wide Webโ?
https://youtube.com/shorts/MxuX7M661Hg
#www #sysadmin #devops #sre #pronunciation #tutorial #software #developers
loading . . .
How do you pronounce โwwwโ the abbreviation for โWorld Wide Webโ? #shorts
#www #sysadmin #sysadminlife #devops #sre #pronunciation #tutorial #software #opensource #developers #techtok
https://youtube.com/shorts/MxuX7M661Hg
about 3 years ago
1
3
3
How do you pronounce "sudo" the #linux/#unix command? So are you team "Su DOUGH" or team "Su DOOOO"
https://www.youtube.com/shorts/qpi5wYblQfY
loading . . .
How do you pronounce "sudo" the #linux/#unix command? #shorts
How do you pronouce "sudo" the #linux/#unix command? #sudo #sysadmin #sysadminlife #devops #sre #pronounciation #tutorial #software #opensource #developers #...
https://www.youtube.com/shorts/qpi5wYblQfY
about 3 years ago
0
3
1
This is probably the most requested pronunciation video i've gotten. How do you say: "fsck" - a.k.a - File System Check. There is no agreed upon pronunciation of this one!
https://youtube.com/shorts/7b-X6MJGkdA
#linux #sysadmin #devops #sre
loading . . .
There is no agreed upon way to pronounce \
#techtips #data #softwareengineer #developer #devops #sysadmin #software #linux #techlife #tutorial #pronounciation #shorts #linux
https://youtube.com/shorts/7b-X6MJGkdA
about 3 years ago
0
1
0
Another episode of โHow do you sayโ. This one is definitely one of my favorites. How do you say JWT (JSON Web Token)?
https://youtube.com/shorts/D2D9umQMKhA?feature=share
loading . . .
Have you heard of JWT before? But how would YOU pronounce it? #shorts #software #howdoyousay
Have you heard of JWT before? But how would YOU pronounce it? #shorts #software #howdoyousay #tech #softwareengineering #softwaretutorials
https://youtube.com/shorts/D2D9umQMKhA?feature=share
about 3 years ago
1
2
2
reposted by
Pete Cheslock
Did you know there are at least 3 (THREE) different ways to say SQL?
https://www.tiktok.com/t/ZTRK9rNSh/
loading . . .
petecheslock on TikTok
Did you know there are 3 ways to pronouce #SQL? #techtok #data #softwareengineer #developer #tech #code
https://www.tiktok.com/t/ZTRK9rNSh/
about 3 years ago
0
1
1
Did you know there are at least 3 (THREE) different ways to say SQL?
https://www.tiktok.com/t/ZTRK9rNSh/
loading . . .
petecheslock on TikTok
Did you know there are 3 ways to pronouce #SQL? #techtok #data #softwareengineer #developer #tech #code
https://www.tiktok.com/t/ZTRK9rNSh/
about 3 years ago
0
1
1
How do you sayโฆ.. โEpochโ Turns out this one was heavily contested on pronounciation.
https://www.tiktok.com/t/ZTRws7M3b/
loading . . .
petecheslock on TikTok
Hey #techtok How do you sayโฆ โepochโ https://en.m.wikipedia.org/wiki/Epoch_(computing) #p#pronouncet#technologys#softwaredevelopers
https://www.tiktok.com/t/ZTRws7M3b/
about 3 years ago
0
0
0
reposted by
Pete Cheslock
SnowSkater for Life
about 3 years ago
AppMap is amazing.
0
1
1
Oooh neat. New handle time.
about 3 years ago
2
1
0
Well - after recording over 30 of these. Check out the promo for my upcoming video series. "How do you say" - Where the words are made up and the pronunciations don't matter. This was an absolute blast and huge thanks to EVERYONE who was able to join me.
https://youtu.be/fkwF2rjOJKI
loading . . .
How Do You Say Tech Lingo - Promo
In the tech community, there are many words and acronyms! This series asks people from different tech companies and communities to pronounce them ๐ฌ Let's see how they sound! ๐ What made you laugh, what made you cringe? Tell us in the comments. We'd love if you try out AppMap as a code editor extension / plugin to visualize your runtime code. We promise it's not like anything else you've tried! https://appmap.io/download
https://youtu.be/fkwF2rjOJKI
about 3 years ago
0
1
0
๐ Hello Friends.
about 3 years ago
3
5
0
you reached the end!!
feeds!
log in