<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
  <channel>
    <title>beengineer500</title>
    <link>https://beengineer500.github.io/</link>
    <description>beengineer500의 개인 기술 메모장. 개발, 운영, 자동화 과정에서 배운 것을 짧고 다시 실행 가능하게 기록합니다.</description>
    <item>
      <title>라우팅 판단을 프록시 밖으로 - llm-d의 CR 구조와 Envoy ext_proc</title>
      <link>https://beengineer500.github.io/posts/2026/llmd-envoy-ext-proc/</link>
      <description>llm-d가 어느 vLLM 파드로 요청을 보낼지 정하는 경로를 Envoy 필터 체인과 ext_proc 설정, 그리고 llm-d가 쓰는 CRD 사이의 의존성까지 따라간 기록</description>
      <pubDate>Sun, 20 Sep 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/llmd-envoy-ext-proc/</guid>
    </item>
    <item>
      <title>AWS Trainium에서 vLLM 서빙 - 첫 배포 8분과 재배포 20초를 가르는 Neuron 컴파일</title>
      <link>https://beengineer500.github.io/posts/2026/aws-trainium-vllm-eks/</link>
      <description>trn1.2xlarge 칩 1개의 NeuronCore 2개가 TP 2가 되고 init 컨테이너 컴파일과 S3 캐시가 파드 기동 시간을 8분에서 20초로 바꾸는 경로를 실제 값으로 따라갑니다</description>
      <pubDate>Sun, 13 Sep 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/aws-trainium-vllm-eks/</guid>
    </item>
    <item>
      <title>BF16이 이긴 콜드 캐시 서빙 비교 - Qwen3-14B AWQ의 더 큰 KV 캐시가 TTFT를 줄이지 못한 이유</title>
      <link>https://beengineer500.github.io/posts/2026/bf16-vs-awq-cold-cache/</link>
      <description>Qwen3-14B를 BF16과 AWQ로 각각 서빙해 콜드 캐시에서 처리량과 TTFT를 재봤습니다</description>
      <pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/bf16-vs-awq-cold-cache/</guid>
    </item>
    <item>
      <title>균등 분산과 캐시 지역성은 다른 목표다 - llm-d prefix-aware 라우팅 실측</title>
      <link>https://beengineer500.github.io/posts/2026/llmd-prefix-aware-routing/</link>
      <description>Kubernetes Service 직결과 load-only 라우터와 prefix-aware 라우터에 같은 프리픽스 반복 워크로드를 흘려 backend request 분포와 prefix-cache hit rate를 비교한 기록</description>
      <pubDate>Sun, 06 Sep 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/llmd-prefix-aware-routing/</guid>
    </item>
    <item>
      <title>최적화 2편. 추측 디코딩과 긴 컨텍스트 - Speculative Decoding과 MEDUSA·EAGLE, Ring Attention</title>
      <link>https://beengineer500.github.io/posts/2026/optimization-02-speculative-decoding-and-long-context/</link>
      <description>작은 초안 모델이 추측한 토큰을 큰 모델이 한 번의 가중치 읽기로 검증하는 추측 디코딩, 그것을 개선한 MEDUSA와 EAGLE, 그리고 KV Cache를 여러 장치에 나눠 담는 Ring Attention을 정리했습니다</description>
      <pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/optimization-02-speculative-decoding-and-long-context/</guid>
    </item>
    <item>
      <title>최적화 3편. 여러 GPU에 모델 나누기 - 통신 패턴과 Tensor·Pipeline Parallel, MoE</title>
      <link>https://beengineer500.github.io/posts/2026/optimization-03-parallelism/</link>
      <description>모델이 GPU 한 장에 안 들어갈 때 여러 장에 나누는 방법과 그때 생기는 통신 병목을 정리했습니다. 집합 통신 다섯 패턴, Tensor·Pipeline Parallel, Prefill-Decode 분리, MoE의 all-to-all 병목까지 다룹니다</description>
      <pubDate>Sun, 30 Aug 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/optimization-03-parallelism/</guid>
    </item>
    <item>
      <title>최적화 1편. 배칭과 어텐션 - GPU 유휴 시간과 KV 캐시를 줄이는 다섯 기법</title>
      <link>https://beengineer500.github.io/posts/2026/optimization-01-batching-and-attention/</link>
      <description>요청을 언제 묶고 언제 쪼갤지 정하는 세 개의 상한, KV 캐시를 줄이는 네 가지 어텐션, 그리고 HBM 왕복을 없애는 커널을 정리했습니다</description>
      <pubDate>Sun, 23 Aug 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/optimization-01-batching-and-attention/</guid>
    </item>
    <item>
      <title>서빙 1편. 여섯 컴포넌트 - 프로세스를 나누는 이유</title>
      <link>https://beengineer500.github.io/posts/2026/serving-01-architecture-and-isolation/</link>
      <description>단일 모델 LLM 서빙 시스템이 왜 여섯 조각으로 갈라지는지, 요청 하나가 프로세스 경계를 넘어 돌아오는 길을 따라갑니다</description>
      <pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/serving-01-architecture-and-isolation/</guid>
    </item>
    <item>
      <title>서빙 2편. 정적 배칭 - 처리량이 64배 벌어지는 지점</title>
      <link>https://beengineer500.github.io/posts/2026/serving-02-static-batching/</link>
      <description>요청 경계를 무시하고 프롬프트를 다시 묶는 배칭의 원리와, 동시 요청을 1개에서 64개까지 올리며 직접 잰 처리량</description>
      <pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/serving-02-static-batching/</guid>
    </item>
    <item>
      <title>서빙 3편. 스트리밍과 vLLM - 303줄을 20줄로</title>
      <link>https://beengineer500.github.io/posts/2026/serving-03-streaming-and-vllm/</link>
      <description>배치 하나에서 나온 토큰을 요청별로 갈라 보내는 SSE 구조와, 직접 만든 서빙 루프를 프레임워크로 갈아탈 때 드러나는 것</description>
      <pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/serving-03-streaming-and-vllm/</guid>
    </item>
    <item>
      <title>서빙 4편. 멀티 모델 - LRU 캐시와 Triton 위임</title>
      <link>https://beengineer500.github.io/posts/2026/serving-04-multi-model-and-triton/</link>
      <description>한 서비스가 여러 모델을 나눠 쓰는 구조와, 무엇을 메모리에 올려둘지 정하는 캐시 정책, 그리고 비용과 지연 사이의 맞바꿈</description>
      <pubDate>Sun, 16 Aug 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/serving-04-multi-model-and-triton/</guid>
    </item>
    <item>
      <title>트랜스포머 1편. 텍스트가 벡터가 되기까지 - 토큰화와 임베딩</title>
      <link>https://beengineer500.github.io/posts/2026/transformer-01-tokenize-embedding/</link>
      <description>GPT-2가 입력 문장을 어떻게 숫자로 바꾸는지, 토큰화부터 768차원 임베딩과 위치 인코딩까지 따라갑니다</description>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/transformer-01-tokenize-embedding/</guid>
    </item>
    <item>
      <title>트랜스포머 2편. 어텐션은 실제로 뭘 계산하나</title>
      <link>https://beengineer500.github.io/posts/2026/transformer-02-attention/</link>
      <description>셀프 어텐션이 Q와 K와 V를 만들고 내적하고 마스킹하고 Softmax를 거쳐 문맥 벡터를 만들기까지, 그리고 헤드를 12개로 쪼개는 이유</description>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/transformer-02-attention/</guid>
    </item>
    <item>
      <title>트랜스포머 3편. 벡터가 다음 단어가 되기까지 - MLP와 샘플링</title>
      <link>https://beengineer500.github.io/posts/2026/transformer-03-mlp-and-sampling/</link>
      <description>어텐션 뒤에서 각 토큰을 혼자 다듬는 MLP, 블록을 지탱하는 세 장치, 그리고 마지막 토큰 하나를 5만 개 단어와 대조해 다음 단어를 고르는 과정</description>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/transformer-03-mlp-and-sampling/</guid>
    </item>
    <item>
      <title>트랜스포머 4편. 서빙에서 어텐션 비용 줄이기 - KV Cache와 prefill과 Flash Attention</title>
      <link>https://beengineer500.github.io/posts/2026/transformer-04-serving/</link>
      <description>토큰 하나 뽑을 때마다 반복되는 어텐션 계산을 KV Cache로 없애고, 첫 계산을 prefill로 나누고, Flash Attention으로 메모리 이동을 줄이는 방법</description>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <guid>https://beengineer500.github.io/posts/2026/transformer-04-serving/</guid>
    </item>
  </channel>
</rss>
