Skip to content

Commit f1d4569

Browse files
kmink3225claude
andcommitted
Add Korean (한국어) pages with EN/KO toggle (English default)
ko_about (/ko/), ko_cv (/ko/cv/), ko_projects (/ko/projects/) + toggle links on EN pages. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1 parent 91ff98e commit f1d4569

6 files changed

Lines changed: 163 additions & 2 deletions

File tree

_pages/about.md

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -26,6 +26,8 @@ latest_posts:
2626
limit: 3
2727
---
2828

29+
<div style="text-align: right; margin-bottom: 1rem;"><strong>English</strong> · <a href="/ko/">한국어</a></div>
30+
2931
I am an **AI Engineer / Data Scientist** with 7+ years of experience. I architect and build **enterprise AI platforms** end-to-end — RAG, LLM agents, and NLP — and I back them with **statistically rigorous evaluation**. My foundation is biostatistics (M.S., Columbia) applied to diagnostics, so I care as much about *measuring* a system as building it.
3032

3133
- **AI/LLM Engineering** — multi-agent RAG platforms (Self-RAG / CRAG / Graph RAG), self-built orchestration harnesses, and LLM-as-judge evaluation pipelines. I build systems that are *grounded and measured*, not demos.

_pages/cv.md

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -4,9 +4,10 @@ permalink: /cv/
44
title: CV
55
nav: true
66
nav_order: 5
7-
cv_pdf: /assets/pdf/example_pdf.pdf # you can also use external links here
87
cv_format: rendercv # options: rendercv, jsonresume
9-
description: This is a description of the page. You can modify it in '_pages/cv.md'. You can also change or remove the top pdf download button.
8+
description: Curriculum vitae — Kwangmin Kim, AI Engineer / Data Scientist.
109
toc:
1110
sidebar: left
1211
---
12+
13+
<div style="text-align: right; margin-bottom: 1rem;"><strong>English</strong> · <a href="/ko/cv/">한국어</a></div>

_pages/ko_about.md

Lines changed: 39 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,39 @@
1+
---
2+
layout: about
3+
title: 소개
4+
permalink: /ko/
5+
nav: false
6+
subtitle: AI Engineer / Data Scientist · 엄밀한 평가로 뒷받침하는 엔터프라이즈 AI 플랫폼(RAG·LLM 에이전트)
7+
8+
profile:
9+
align: right
10+
image: prof_pic.png
11+
image_circular: true
12+
more_info: >
13+
<p>대한민국, 서울</p>
14+
<p><a href="mailto:kmink3225@gmail.com">kmink3225@gmail.com</a></p>
15+
16+
selected_papers: false
17+
social: true
18+
19+
announcements:
20+
enabled: false
21+
scrollable: true
22+
limit: 5
23+
24+
latest_posts:
25+
enabled: false
26+
scrollable: true
27+
limit: 3
28+
---
29+
30+
<div style="text-align: right; margin-bottom: 1rem;"><a href="/">English</a> · <strong>한국어</strong></div>
31+
32+
7년 이상 경력의 **AI Engineer / Data Scientist** 다. RAG, LLM 에이전트, NLP 기반의 **엔터프라이즈 AI 플랫폼을 아키텍처부터 설계·구축**하고, 이를 **통계적으로 엄밀한 평가**로 뒷받침한다. 진단 분야에 적용된 생물통계(컬럼비아 석사)가 토대라, 시스템을 *만드는 것*만큼 *측정하는 것*을 중요하게 여긴다.
33+
34+
- **AI/LLM 엔지니어링** — 멀티 에이전트 RAG 플랫폼(Self-RAG / CRAG / Graph RAG), 자체 오케스트레이션 하네스, LLM-as-judge 평가 파이프라인. 데모가 아니라 *접지(grounding)되고 측정되는* 시스템을 만든다.
35+
- **데이터 사이언스** — 실험설계, 인과추론, 생존분석, 딥러닝/NLP. 이론적 배경은 [1,900편 이상의 기술 블로그 글](https://kk3225.netlify.app)로 정리해 왔다.
36+
37+
특허 **7건**(제1발명가 4건)을 출원했고, R&D 부문 President's Award(Seegene)와 생물통계 Chair's Award(컬럼비아)를 받았다.
38+
39+
자세한 내용은 [이력서](/ko/cv/)[프로젝트](/ko/projects/)를, 깊이는 [블로그](https://kk3225.netlify.app)를, 연락은 [GitHub](https://github.com/kmink3225) · [LinkedIn](https://www.linkedin.com/in/kwangmin-kim-a5241b200/)을 참고하면 된다.

_pages/ko_cv.md

Lines changed: 71 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,71 @@
1+
---
2+
layout: page
3+
title: 이력서
4+
permalink: /ko/cv/
5+
nav: false
6+
description: 김광민 — AI Engineer / Data Scientist 이력
7+
toc:
8+
sidebar: left
9+
---
10+
11+
<div style="text-align: right; margin-bottom: 1rem;"><a href="/cv/">English</a> · <strong>한국어</strong></div>
12+
13+
**김광민 (Kwangmin Kim)** · AI Engineer / Data Scientist · 대한민국 서울
14+
[kmink3225@gmail.com](mailto:kmink3225@gmail.com) · [GitHub](https://github.com/kmink3225) · [LinkedIn](https://www.linkedin.com/in/kwangmin-kim-a5241b200/)
15+
16+
## 요약
17+
18+
7년 이상 경력의 AI Engineer / Data Scientist. RAG, LLM 에이전트, NLP 기반 엔터프라이즈 AI 플랫폼을 아키텍처부터 설계·구축하고, 통계적으로 엄밀한 평가로 뒷받침한다. 전문 분야는 LLM 에이전트, RAG/Graph RAG, 딥러닝/NLP, 머신러닝, 실험설계, 통계분석이다.
19+
20+
## 경력
21+
22+
### Seegene — Data Scientist / AI Engineer
23+
24+
*2020.12 – 현재 · 대한민국*
25+
26+
전사 멀티 에이전트 플랫폼의 기술 리드/아키텍트로, 엔터프라이즈 AI 에이전트·RAG 플랫폼과 그 평가 체계를 설계·구축한다. 이전에는 진단 분야의 ML·통계 모델링을 담당했다.
27+
28+
- 도메인 특화 **멀티 에이전트 RAG 지식 플랫폼**을 아키텍처부터 총괄 설계·구축하고, 단일 에이전트 파일럿에서 전사 과제로 확장(초기 ~30명 배포, 전사 확장 중).
29+
- **지식 QnA 챗봇** 설계 — 9개 sub-agent **Self-RAG/CRAG** 루프 + 토큰 스트리밍 + 출처 인용. 151건 질의에서 10개 지표 전수 통과(사용자 만족도 ~98%, 평균 응답 4.66초, 인용률 96.9%, 시스템 성공률 100%), 4모델 LLM-as-judge 평가에서 사실·추론 5.0/5.0.
30+
- **자체 에이전트 오케스트레이션**이 범용 CLI 대비 **건당 비용 최대 ~17배 절감**을 벤치마크로 입증(paired t-test/McNemar/bootstrap CI, 6지표 Composite).
31+
- **NLP 기반 데이터 표준화 시스템** — 검증 시간 **8시간 → 0.73초(99% 단축)**, 메타데이터 일관성 8.4% → 98.7%. 8개 모델 벤치마크(14클래스, 95% CI, McNemar+Holm)에서 KLUE-RoBERTa 96.88% 선정, 671K BiLSTM이 110M 모델과 통계적 동급임을 입증.
32+
- 하드코딩 PCR 신호 baseline 알고리즘을 **데이터 기반 모델**로 재설계, 위음성률 **0.47% → 0.04%(91.49% 개선)**.
33+
- 진단 장비 QC를 **2단계 LSTM + 10개 지표**로 자동화(신호 61,248개), QC 시간 ~93% 단축(연간 운영비 약 13배 절감) — R&D President's Award.
34+
- **모델 평가·MLOps 체계** 구축 — LLM-as-judge 자동 채점 + 아키텍처 A/B 벤치마크 + 메트릭 로깅. IT/BT 20여 명 멘토링.
35+
36+
### 컬럼비아 의대 (CUIMC) — Taub Institute · Statistical Research Assistant
37+
38+
*2018.12 – 2020.05 · 미국 뉴욕*
39+
40+
알츠하이머병 바이오마커 발견을 위한 대규모 다중 오믹스 분석.
41+
42+
- 유전체·대사체·임상데이터 통합으로 약 3,000개 대사물질 중 **핵심 바이오마커 13개(p<0.01)** 발견, 8개월간 미발견 교란자 규명.
43+
- 146 샘플 × 3,000 고차원 변수 처리, 10+ ML 알고리즘 비교 후 sPLS 선정(분류 정확도 84% + 해석력), Cox/GEE로 20년 발병 위험 모델 구축.
44+
45+
## 학력
46+
47+
- **M.S. Biostatistics**, Columbia University (2017–2019) — Chair's Award
48+
- **B.A. Mathematics**, Baruch College, CUNY (2015–2017)
49+
- **B.S. Biochemistry**, 강원대학교 (2006–2012) — 수석 졸업, 학장상
50+
51+
## 기술
52+
53+
- **LLM 에이전트 / GenAI** — RAG, Agentic RAG, Graph RAG, Self-RAG/CRAG, LangChain, LangGraph, Azure OpenAI, Azure AI Search, OpenAI/Claude API, 프롬프트 엔지니어링
54+
- **NLP / 딥러닝** — KLUE-RoBERTa, KoBERT, ALBERT, BiLSTM/LSTM, Hugging Face Transformers, PyTorch, KiwiPiePy/KoNLPy
55+
- **머신러닝 / 통계** — scikit-learn, HDBSCAN, 회귀/생존분석, 시계열, 인과추론(A/B Test), 실험설계
56+
- **데이터 / 백엔드 엔지니어링** — Python, R, SQL(PostgreSQL), SAS, FastAPI, Streamlit, Apache Airflow, Parquet, AST, Docker, Azure DevOps
57+
58+
## 수상 · 특허
59+
60+
- **특허 출원 7건**(제1발명가 4건) — Ct 기반 맞춤 치료법, 의료 플랫폼 구독 시스템, 진단 장비 Noise Test 자동화, 의료 장비 Noise Level 측정 알고리즘 (제1발명가); 분자진단 예측 모델 등 (제2발명가). Seegene, 2021–2022.
61+
- **President's Award** (R&D 부문 우수상), Seegene, 2021
62+
- **Chair's Award** — Graduation Practicum Research Competition, Columbia Biostatistics, 2019
63+
64+
## 자격
65+
66+
- Microsoft Azure — DP-203(Data Engineering), DP-100(Data Science), DP-300(Database), 2025
67+
- SAS Certified Base Programmer, 2018
68+
69+
## 언어
70+
71+
- 한국어 (모국어) · 영어 (유창)

_pages/ko_projects.md

Lines changed: 46 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,46 @@
1+
---
2+
layout: page
3+
title: 프로젝트
4+
permalink: /ko/projects/
5+
nav: false
6+
description: AI/LLM 엔지니어링과 데이터 사이언스 주요 작업. 운영 세부사항은 회사 IP 보호를 위해 추상화했다.
7+
toc:
8+
sidebar: left
9+
---
10+
11+
<div style="text-align: right; margin-bottom: 1rem;"><a href="/projects/">English</a> · <strong>한국어</strong></div>
12+
13+
> 아키텍처와 방법론은 상위 수준으로만 기술한다. 운영 코드와 내부 데이터는 비공개다.
14+
15+
## 엔터프라이즈 멀티 에이전트 RAG 플랫폼
16+
17+
*역할: 기술 리드 / 아키텍트 · 스택: Python, LangChain, LangGraph, Azure OpenAI, Azure AI Search, FastAPI*
18+
19+
도메인 특화 **멀티 에이전트 RAG 지식 플랫폼**을 아키텍처부터 총괄 설계·구축하고, 단일 에이전트 파일럿에서 전사 과제로 확장했다(초기 ~30명 배포, 전사 확장 중).
20+
21+
- **지식 QnA 챗봇** — 9개 sub-agent Self-RAG/CRAG 루프 + 토큰 스트리밍 + 출처 인용. 151건 질의 10개 지표 전수 통과(만족도 ~98%, 4.66초, 인용률 96.9%, 성공률 100%), 4모델 LLM-as-judge에서 사실·추론 5.0/5.0.
22+
- **자체 오케스트레이션 vs 범용 CLI** — 최상위 구성에서 건당 비용 **최대 ~17배 절감**(paired t-test/McNemar/Cohen's d/bootstrap CI, 6지표 Composite).
23+
- **RAG 파이프라인** — Parent-Child + Contextual Chunking, 하이브리드 검색(BM25+Vector), Child→Parent 매핑, 리랭킹으로 환각 억제. LangChain → LangGraph → Agentic 3단계 로드맵.
24+
25+
## NLP 기반 데이터 표준화 시스템
26+
27+
*역할: 기술 리드(IT/BT 20여 명 멘토링) · 스택: PyTorch, Transformers, KLUE-RoBERTa, BiLSTM, HDBSCAN, RAG, pytest, Docker*
28+
29+
현업 메타데이터 불일치 문제를 정의하고 Rule + 분류기 + RAG 하이브리드 시스템으로 해결, 파일럿 성공 후 전사 적용·후속 AI 플랫폼의 출발점이 되었다.
30+
31+
- 검증 시간 **8시간 → 0.73초(99% 단축)**, 부서 간 문의 월 70 → 4건(94.3%↓), 일관성 8.4% → 98.7%, 완전성 29.6% → 100%.
32+
- **8개 모델 벤치마크**(14클래스, 7,698건, stratified, 95% CI, McNemar+Holm) — KLUE-RoBERTa 96.88% 최고, 5-fold CV로 671K BiLSTM이 110M 모델과 통계적 동급(p=0.73) 입증, 1.48ms 경량 배포안 도출.
33+
34+
## 진단 신호 모델링 & QC 자동화
35+
36+
*역할: 프로젝트 PM / Data Scientist · 스택: Python, R, LSTM, 기저함수 모델링, PCA/t-SNE/DBSCAN, R Shiny*
37+
38+
- **PCR 신호 baseline 보정** — 하드코딩 레거시를 데이터 기반 혼합 기저함수 모델로 재설계, 위음성률 **0.47% → 0.04%(91.49% 개선)**, 5종 경쟁 알고리즘 대비 백색잡음 근사도 1위.
39+
- **장비 QC 자동화** — 2단계 LSTM + 10개 지표(장비 2,201대·신호 61,248개), QC 시간 ~93% 단축(연간 운영비 약 13배 절감), 합/불 분류 94.5%. R&D President's Award 및 제1발명가 특허 2건.
40+
41+
## 알츠하이머 멀티 오믹스 바이오마커 발견
42+
43+
*역할: Statistical Research Assistant, 컬럼비아 의대 Taub Institute · 스택: R, Cox/GEE, sPLS, Lasso/Ridge/RF/SVM/GBM*
44+
45+
- 유전체·대사체·임상데이터 통합으로 약 3,000개 대사물질 중 **핵심 바이오마커 13개(p<0.01)** 발견, 8개월간 미발견 교란자 규명.
46+
- 고차원 소표본(146 샘플 × 3,000 변수), 10+ ML 비교 후 sPLS 선정(정확도 84% + 해석력), Cox/GEE로 20년 발병 위험 모델 구축. 연구 경진대회 top-3, Chair's Award, Taub Institute 정규직 제안.

_pages/projects.md

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -9,6 +9,8 @@ display_categories: [work]
99
horizontal: false
1010
---
1111

12+
<div style="text-align: right; margin-bottom: 1rem;"><strong>English</strong> · <a href="/ko/projects/">한국어</a></div>
13+
1214
<!-- pages/projects.md -->
1315
<div class="projects">
1416
{% if site.enable_project_categories and page.display_categories %}

0 commit comments

Comments
 (0)