Conversation
같은 유저의 중복 출근 요청을 막기 위해 attendance:lock:{userId}:{workDate}
분산락(Redisson)을 1차 방어선으로, (user_id, work_date) 유니크 제약을
2차 방어선으로 이중 구조를 적용. BusinessException을 실제 HTTP 상태 코드로
변환하는 GlobalExceptionHandler가 없어 전부 500으로 나가던 버그와,
clockin()이 읽기 전용 트랜잭션에 묶여 INSERT가 실패하던 버그를 함께 수정.
근무조(WorkShift)별 시작 시각 기준 지각 판정 로직 추가.
ADR-0001: 출근 체크 동시성 제어 설계 및 k6 실측 결과 ADR-0002: 성능 개선 후보 검토(비동기 병렬 호출, N+1, GC, 캐시) ADR-0003: 다국어 작업일지 검색 확장 로드맵 ADR-0004: 출퇴근 버스트 트래픽 대응 로드맵 ADR-0005: 근태 시스템 포트폴리오 요소 체크리스트 대비 격차 진단
Docker 빌드/EC2 배포 파이프라인이 push마다 실패해 일단 비활성화. 필요 시 git history에서 복구.
springdoc-openapi-starter-webmvc-ui:2.6.0이 Spring Boot 4.0.1(Spring Framework 7)의 바뀐 ControllerAdviceBean 생성자 시그니처와 호환되지 않아 /v3/api-docs, /swagger-ui.html이 500(NoSuchMethodError)을 내던 문제를 수정. 실제 API 엔드포인트는 영향 없었음.
schema.sql에 스키마만 있고 Java 코드가 없던 두 도메인을 채움 (docs/adr/0005 진단 기준 우선순위). 체크리스트(com.DOCKin.checklist): 장비 작업 전/후 점검. checklist_results를 항목(item) 단위 append-only 감사 로그로 재설계(구 스키마는 체크리스트 통짜로 is_checked 하나만 둬서 부분 완료를 표현 못 했음). 현재 상태 조회는 항목별 최신 결과를 배치 쿼리 1번으로 병합해 N+1을 피함. 휴가(com.DOCKin.absence): 연차/병가 신청-승인 워크플로우. 승인 시점에만 잔여 연차(Member.remainingLeaveDays)를 검증/차감해 여러 PENDING 요청의 초과 예약을 방지. absence_requests의 미사용 컬럼(last_message_content/at, chat_rooms에서 복사된 것으로 보임)을 승인/거절 사유 코멘트(decision_comment)로 재활용. 두 도메인 모두 기존 관례(서비스 레이어 수동 RBAC, BusinessException/ErrorCode)를 그대로 따름. 단위 테스트 30개(체크리스트 20 + 휴가 10) 추가, 로컬 Docker 환경에서 curl로 전체 플로우 실기동 검증 완료.
PROJECT-SCOPE.md: DOCKin이 팀플이고 본인 담당이 Spring+DB(FastAPI는 팀원이 번역/STT만 담당, 임베딩 검색 없음)임을 명시 - 포폴 서술 시 담당 범위를 벗어난 내용을 본인 것처럼 쓰지 않기 위함.
챗봇이 사용자 질문을 FastAPI로 그대로 넘기기만 해 근거 없이 답하던 구조에, Spring 쪽에 검색 단계를 넣었다. 생성(G)은 기존 FastAPI가 그대로 담당한다. ## 구성 - 임베딩 추론 서버를 별도 컨테이너로 분리 (TEI / multilingual-e5-small) JVM 힙 400M에 모델 적재가 불가하고, 분리하면 모델 교체가 배포와 무관해진다 - document_chunks: 작업일지 / 번역본 / 안전교육을 청킹해 384차원 벡터로 적재 - 브루트포스(정확 최근접) 검색. 챗봇이 약 0.02 TPS라 지연 요구가 없고, recall이 100%라 Phase 2에서 ANN 도입 시 정답 기준선이 된다 - 팀원 FastAPI의 /api/chatbot 계약은 바꾸지 않았다. 근거를 별도 필드가 아니라 user 메시지 content 안에 녹여 보낸다 ## 권한 - 후필터가 아니라 선필터 벡터 유사도만으로는 접근 제어를 할 수 없다. top-k를 뽑고 거르면 요청한 k보다 적게 반환되어 "내가 못 보는 문서가 있다"는 사실이 유출되고, 검색어를 바꿔가며 반복하면 문서의 존재 윤곽을 그릴 수 있다. SQL 단계에서 후보에 아예 넣지 않는다. 임베딩 장애 시 키워드 폴백도 같은 선필터를 쓴다 - 장애 때만 권한이 느슨해지면 안 된다. ## 실측 (docs/SERVICE-SCALE-ASSUMPTIONS.md) - 임베딩 배치 12ms/건 (단건 99ms 대비 8배) -> 인덱싱은 반드시 배치 - 브루트포스 10만 청크: 일반 사용자 1,485ms / 관리자 전체 5,165ms - EXPLAIN: 권한 조건 OR가 인덱스를 못 타 10만 행을 읽고 90%를 버리고 있었다 -> 두 조회로 분리해 770 -> 659ms. 다만 14%에 그친다 - 진짜 병목은 DB가 아니라 전송이었다. 서버 실행 925ms vs JDBC 왕복 4,882ms로 81%가 146MB를 애플리케이션으로 가져오는 비용이다. 쿼리 튜닝으로는 못 건드린다 - 교차언어: 영어 recall@1 5/5, 베트남어 4/5 (표본 5쌍) 절대 유사도는 오답도 0.800이라 임계값 기준으로 쓸 수 없다 -> min-score 비활성 ## 함께 정리한 것 - ChatHistory 삭제: 참조가 한 곳도 없으면서 ChatLog와 같은 테이블에 중복 매핑 - TranslateLog를 work_log_translations로 통일 (엔티티만 translate_logs를 보고 있었다) UNIQUE(log_id, language_code) 추가 - 재번역 시 중복 행이 쌓여 색인을 오염시켰다 - 인덱싱 순회를 OFFSET에서 커서(keyset)로. 성능뿐 아니라 순회 중 INSERT 시 행이 누락되는 문제가 있었다 - schema.sql의 work_log_translations 중복 컬럼 수정 (스크립트가 실행 불가 상태였다) ## 알려진 한계 - JDBC 배치 INSERT가 동작하지 않는다. IDENTITY 전략이 막고 있고 설정은 조용히 무시된다(Com_insert로 확인, 10,000건 저장에 문장 10,000개). JdbcTemplate으로 우회 가능하나 PostgreSQL 이전 후 SEQUENCE로 근본 해결되므로 임시 부채를 만들지 않았다. 검증 테스트를 남겨 이전 후 같은 방법으로 확인한다 - 리랭커/하이브리드 검색/Redis 캐시는 근거가 확보되면 도입한다 (ADR-0006 10절) ErrorCode.java에는 진행 중인 근태 작업의 에러 코드 2개가 함께 담겼다. 같은 파일이라 분리할 수 없었고, enum 상수 추가라 동작에 영향은 없다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
휴가 도메인과 근태 도메인이 각각 완성되어 있는데 사이가 비어 있었다. AbsenceRequestService.approveRequest()는 상태만 APPROVED로 바꾸고 끝났고, absence 패키지에는 Attendance 참조가 한 곳도 없었다. 그 결과 AttendanceStatus.VACATION은 정의만 되고 채우는 곳이 없었으며, 승인된 휴가일이 결근 처리 대상으로 남아 있었다. ## 휴가 승인 -> 근태 반영 - AbsenceApprovedEvent 발행 / AbsenceApprovedListener 수신 - 이벤트로 분리하되 트랜잭션은 분리하지 않았다. 비동기나 AFTER_COMMIT으로 두면 근태 반영이 실패했을 때 승인만 남아 "승인된 휴가인데 결근" 상태가 된다. 동기 리스너라 같은 트랜잭션에서 실행되고 실패 시 함께 롤백된다 - 소급 승인 시 기존 기록은 덮어쓰지 않는다. 이미 출근한 날의 기록을 지우면 근무한 사실이 사라진다. 덮어쓰기는 관리자의 명시적 수정으로 다룰 문제다 - 기간 전체를 한 번에 조회한다. 날짜별로 조회하면 기간만큼 쿼리가 늘어난다 ## 자정 결근 배치 - 전날 기준 미체크 인원을 ABSENT로 생성. 휴가 기록이 있는 사람은 조회 단계에서 제외된다 - Clock을 주입받는다. LocalDate.now()를 직접 부르면 "어제"가 실행 시각에 따라 달라져 테스트가 불가능하다. 근태는 날짜 경계가 곧 비즈니스 규칙이다 - @scheduled 메서드에 @transactional을 직접 붙였다. 내부 호출 대상에만 붙이면 자기호출이라 프록시를 거치지 않아 트랜잭션이 걸리지 않는다 ## 스키마 Attendance.clockInTime을 nullable로 바꿨다. NOT NULL은 "출근한 날"만 상정한 제약이라 휴가·결근을 담을 수 없었다. 더미 시각을 넣으면 "0시에 출근한 기록"이 되어 근무시간 집계를 오염시키므로 없는 것은 없는 대로 둔다. ## 알려진 한계 - 공휴일 근무일 캘린더가 없어 주말만 제외한다. 이대로 운영에 넣으면 공휴일에 전원이 결근 처리된다. WorkShift는 교대 시간대만 정의하고 휴무일 정보가 어디에도 없다. ADR-0005의 "근무 정책 엔진 미충족"과 같은 뿌리이며 P2-6으로 남겼다. ## 테스트 AbsenceApprovedListenerTest 6건 / AttendanceBatchServiceTest 5건 추가. AbsenceRequestServiceTest에 ApplicationEventPublisher 목을 추가했다 - 생성자 의존성이 늘어 기존 승인 테스트 2건이 NPE로 깨졌던 것을 고친 것이다. ADR-0005 기준 휴가 연동 / 상태값 활용 / 배치 / 이벤트 기반 설계 네 항목이 닫혔다. 연차 차감(P2-4)은 이미 구현되어 있었다 - 백로그 기재가 오기였다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
결근 배치가 주말만 제외하고 공휴일을 판단하지 못했다. 이대로 운영에 넣으면 공휴일에 전원이 결근 처리된다. WorkShift는 교대 시간대만 정의하고 휴무일 정보가 어디에도 없었다. ## 설계 - 날짜를 PK로 쓰는 자연키. 같은 날이 두 번 등록될 수 없어야 하는데 대리키 + 유니크 제약보다 자연키가 의도를 직접 드러낸다 - 등록되지 않은 날은 기본 규칙(평일=근무, 주말=휴무)을 따른다. 캘린더를 비워둬도 기존 동작이 유지되어 점진적으로 채울 수 있다. 반대로 "미등록=휴무"로 잡으면 캘린더를 채우기 전까지 배치가 조용히 무력화된다 - 양방향 예외를 표현한다. 평일인데 쉬는 날(공휴일)과 주말인데 일하는 날(특근)이 둘 다 존재하므로, DayType.WORKDAY를 주말에 등록하면 특근일이 된다 - 근무일 판단을 WorkCalendarService 한 곳으로 모았다. 초과근무 계산과 월말 집계도 같은 판단을 필요로 하게 되므로 각자 요일을 보게 두면 곧 어긋난다 ## 남은 것 등록 수단이 없다. register()/registerAll()은 있으나 HTTP 엔드포인트가 없어 현재는 SQL로만 넣을 수 있다. 관리자 컨트롤러와, 공공데이터 API로 법정공휴일을 연 1회 자동 적재하는 연동이 뒤따라야 한다. 이 캘린더는 전사 공통 휴무일만 다룬다. 조선소는 교대조마다 휴무 패턴이 달라 실제로는 (날짜, 교대조) 단위로 근무일이 결정되지만 그건 근무 정책 엔진의 영역이다 (ADR-0005 "근무 정책 엔진 미충족", 백로그 P3). ## 테스트 WorkCalendarServiceTest 7건 추가 (미등록 평일/주말, 평일 공휴일, 주말 특근, 회사 휴무일, 중복 등록 시 갱신, 권한 검증). AttendanceBatchServiceTest에 평일 공휴일 케이스를 추가하고 새 의존성을 반영했다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
LocalDate.now()/LocalDateTime.now()를 직접 호출하면 시간 의존 로직을 결정론적으로 테스트할 수 없다. Clock을 주입받아 테스트에서 Clock.fixed(...)로 대체 가능하게 한다. 근태는 날짜 경계가 곧 비즈니스 규칙이라(지각 판정, "어제" 기준 결근 처리) 특히 중요하다. AttendanceBatchService가 이 빈을 필요로 하므로 먼저 커밋한다 - 없으면 기동에 실패한다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
## Clock 주입
LocalDate.now()/LocalDateTime.now()를 직접 호출하던 지점을 Clock 기반으로 바꿨다.
지각 판정(교대 시작 시각 비교)과 날짜 기준 조회가 실행 시각에 따라 달라져
결정론적으로 테스트할 수 없었다. 근태는 날짜·시각 경계가 곧 비즈니스 규칙이라
테스트에서 Clock.fixed(...)로 고정할 수 있어야 한다.
## 퇴근 로직 예외 정리
- 출근 기록이 없을 때 USER_NOT_FOUND를 던지고 있었다.
사용자는 존재하는데 오늘 출근을 안 한 상황이므로 ATTENDANCE_NOT_CHECKED_IN으로 바꿨다
- 이미 퇴근한 경우 IllegalArgumentException("출근 기록이 존재하지 않습니다")를 던졌다.
도메인 예외가 아니었고 메시지도 실제 상황과 반대였다.
ATTENDANCE_ALREADY_CHECKED_OUT으로 바꿨다
## 버그 수정
setClockOutTime(LocalDateTime.now())가 now 변수와 별도로 다시 현재 시각을 호출하고 있었다.
근무시간 계산에 쓰는 now와 실제 저장되는 clock_out_time이 미세하게 어긋났다.
now 변수를 그대로 쓰도록 고쳤다.
## 테스트
AttendanceServiceTest 9건 추가. Clock 고정으로 지각/정상 판정 경계를 검증한다.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
saveTranslateLog()의 Mono.zip 병렬화에 대한 벤치마크와 그 결과를 남긴다. JDK HttpServer로 300ms 인위 지연 스텁 서버를 띄워 측정했다. | 방식 | 소요 시간(워밍업 후) | |---|---| | 순차 .block() 2회 | 622ms | | Mono.zip() 병렬화 | 320ms | 약 1.9배로, 300ms 지연 두 개를 겹치느냐 아니냐라는 이론값과 일치한다. ## 측정 과정에서 걸러낸 두 가지 착시 - 첫 측정은 순차 2540ms vs 병렬 613ms(4.1배)였으나, 대부분 Reactor Netty 커넥션 초기화가 순차 버전에서 두 번 발생한 콜드 스타트 효과였다. 워밍업 호출 5회를 추가하니 1.9배로 수렴했다 - 스텁 서버에 setExecutor()를 지정하지 않아 요청을 한 번에 하나씩만 처리하고 있었다. 이 상태에서는 클라이언트가 병렬로 보내도 서버에서 직렬화되어 효과가 사라진다. Executors.newFixedThreadPool(4)로 고친 뒤에야 위 수치가 나왔다 측정 조건을 명시해 둔다 - 로컬 스텁 서버 기준이며 실제 FastAPI 환경에서는 다를 수 있다 (ADR-0001과 동일한 관례). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
approveRequest()가 잔여 연차를 읽고-검사하고-쓰는데 아무 동시성 제어가 없었다. 관리자 두 명이 같은 사용자의 신청 두 건을 동시에 승인하면 둘 다 낡은 잔액을 읽어 둘 다 검사를 통과하고 둘 다 차감한다. ## 실측 (LeaveBalanceConcurrencyTest) 잔액 5일에 3일짜리 신청 두 건 동시 승인. "막힌다"만 보이면 애초에 문제가 있었는지 알 수 없어 락 없는 경우도 함께 측정했다. | 조건 | 승인 성공 | 소비된 연차 | DB 잔액 | |---|---|---|---| | 락 없음 | 2건 | 6일 | 2일 | | SELECT ... FOR UPDATE | 1건 | 3일 | 2일 | 락이 없으면 잔액 5일을 초과해 6일이 승인되고 3일치가 증발한다. ## 왜 ADR-0001의 결론을 그대로 쓸 수 없나 출근 체크는 "같은 행이 두 번 생성"되는 문제라 DB 유니크 제약이 최종 방어선이 됐다. 연차는 "수치 갱신이 덮어써지는" 문제라 제약으로 표현할 수 없고 락이 유일한 수단이다. "DB가 최종 방어선"이라는 사고방식은 유지되지만 수단이 제약에서 락으로 바뀐다. 분산락을 쓰지 않은 이유는 요청 패턴이다. 출근은 피크에 몰려 DB 커넥션을 오래 잡지 않으려고 Redis를 앞단에 뒀지만, 휴가 승인은 경합이 드물고 단일 DB라 저장소를 하나 더 끌어들일 이유가 없다. ## 검토했으나 택하지 않은 대안 - 낙관적 락(@Version): 충돌 시 재시도 로직이 필요한데 경합이 드물어 복잡도만 는다 - 원자적 UPDATE(SET remaining = remaining - :days WHERE remaining >= :days): 읽기 단계가 없어 lost update가 성립하지 않는 가장 가벼운 해법이지만, 벌크 UPDATE라 영속성 컨텍스트와 어긋난다. 승인 로직이 이미 엔티티를 다루고 있다 ## 함께 넣은 방어 Member.useLeaveDays()가 잔액 부족 시 예외를 던진다. 불변식을 데이터 곁에 두어 어떤 경로로 들어와도 음수가 되지 않게 한다. 다만 이것만으로 동시성은 해결되지 않는다 - 두 트랜잭션이 각각 낡은 값을 읽으면 둘 다 통과한다. ADR-0001에 4-1절로 기록했다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
비관적 락의 대기 한도가 방치되어 MySQL 기본값 50초가 그대로 적용되고 있었다. 관리자가 승인 버튼을 누르고 50초를 기다린 끝에 실패를 보는 것은 최악의 경험이고, 그동안 DB 커넥션도 계속 묶인다. ## 먼저 시도했다가 버린 방법 - JPA 힌트 jakarta.persistence.lock.timeout을 3초로 걸었으나 조용히 무시됐다. 힌트 설정값 : 3,000 ms 실제 대기 시간 : 50,850 ms 결과 : CannotAcquireLockException 생성된 SQL에도 대기 시간이 없었다: select ... from users m1_0 where m1_0.user_id=? for update of m1_0 표준 JPA 힌트지만 DB가 문장 단위 대기 시간을 지원해야 적용된다. MySQL은 NOWAIT과 SKIP LOCKED만 지원하고 임의의 초를 받는 문법이 없다 (Oracle의 FOR UPDATE WAIT n에 해당하는 것이 없다). 앞서 hibernate.jdbc.batch_size가 IDENTITY 전략 때문에 무시되던 것과 같은 종류의 함정이다. 설정을 걸었다는 사실과 그것이 동작한다는 사실은 다르다. ## 채택한 방법 - 서버 파라미터 compose.yaml에 --innodb-lock-wait-timeout=5 추가. 재측정 결과 50,850ms -> 5,878ms로 적용을 확인했다. 남은 한계: MySQL은 세션 단위까지만 조정할 수 있어 서버 전역 5초로 고정된다. 인덱싱 배치처럼 오래 잡아도 되는 작업에는 짧을 수 있다. PostgreSQL은 SET LOCAL lock_timeout으로 트랜잭션 단위 조정이 가능해 Phase 2에서 이 트레이드오프가 사라진다. ## 문서 WORK-BACKLOG P1-11에 "이전 시 함께 검증할 항목" 표를 추가했다. MySQL에서 발견한 문제들이 PostgreSQL에서 어떻게 되는지 정리하되, 예상은 예상으로 표시했다 - MySQL에서 세 번 다 예상이 빗나갔다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ADR-0006이 계획한 Phase 2의 앞 절반. pgvector 도입(2b) 전에 스택 이관만 먼저 하고 기능은 그대로 둬서, 변경이 겹치지 않게 했다. ## 이관 근거 벡터 검색에 ANN 인덱스가 필요한데 MySQL 8.0에 없다(9.x의 VECTOR도 인덱스는 HeatWave 전용). 브루트포스 10만 청크가 5.2초이고 병목의 81%가 벡터를 애플리케이션으로 가져오는 전송이라, 유사도 계산을 DB 안에서 끝내는 pgvector가 유일한 해법이다. 트리거를 "청크 10만 초과"로 적어뒀으나 실사용 트래픽이 없어 도달 불가능한 조건이었다. 실제 근거는 "합성 데이터로 한계를 실측했고, 이관할 데이터가 없는 지금이 가장 싸다"이다. ## 변경 - 드라이버/URL/방언 전환. 이미지는 pgvector 포함본(2b에서 다시 안 바꾸도록) - DocumentChunk, Attendance를 IDENTITY -> SEQUENCE(allocationSize=50). 대량 적재 대상이라 IDENTITY가 JDBC 배치를 막고 있었다 - DocumentChunk.embedding: VARBINARY(4096) -> BYTEA - lock_timeout=5s, pg_stat_statements 활성화 - 검증 테스트 4개를 PostgreSQL로 이관 ## 이관 중 잡은 문제 SafetyCourse.createdAt에 columnDefinition = "DATETIME"이 하드코딩돼 있어 type "datetime" does not exist로 테이블 생성이 실패했다. ddl-auto=update가 DDL 오류를 로그만 남기고 기동을 막지 않아 safety_courses가 없는 채로 앱이 정상 기동한 것처럼 보였다. columnDefinition 전수 조사로 잡았고, 타입 명시를 제거해 방언이 고르게 했다. ## 재측정 (SERVICE-SCALE-ASSUMPTIONS 6-5) 브루트포스 검색 — 코드 동일, DB만 교체 | 적재 | 경로 | MySQL | PostgreSQL | |---|---|---|---| | 100,000 | 선필터 | 1,485ms | 424ms | | 100,000 | 전체 스캔 | 5,165ms | 4,059ms | 전체 스캔은 비슷한데 선필터만 3.5배 빨라졌다. EXPLAIN으로 원인을 확인했다 - MySQL이 못 골랐던 계획을 PostgreSQL은 BitmapOr로 처리한다(770ms -> 41.5ms). 같은 인덱스를 두 번 스캔해 비트맵으로 OR 연산한 뒤 힙에 접근한다. 대량 적재 — SEQUENCE 전환으로 nextval 200회(increment_by=50, 10,000건 적재). IDENTITY였다면 건별 왕복 10,000회였을 자리다. 지표를 문장 수에서 시퀀스 호출 수로 바꿔야 했다 - PostgreSQL 드라이버는 배치를 파이프라인으로 실행해 문장 수로는 배치 여부를 구분할 수 없다. 락 대기 — 서버 설정 5,171ms, SET LOCAL 지정 시 1,004ms. MySQL에서는 서버 전역으로만 조정할 수 있어 "승인에는 적당하지만 배치에는 짧다"는 문제가 남았는데, 트랜잭션 단위 조정이 되면서 사라졌다. ## 유지한 결정 권한 OR 조건의 두 쿼리 분리를 되돌리지 않는다. BitmapOr가 확인됐지만 단일 쿼리 41.5ms 대 분리 18.5ms로 분리가 여전히 조금 빠르다. 전체 테스트 83건 통과(MySQL 시절 DB 미연결로 늘 실패하던 contextLoads 포함). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ADR-0006 Phase 2b의 첫 단계. 컬럼 타입만 바꾸고 검색 쿼리는 그대로 두었다.
따라서 이 커밋만으로는 성능이 달라지지 않는다 - 인덱스와 쿼리 이전의 토대다.
## 전제가 틀렸다
ADR-0006 8절과 백로그는 이 작업을 "Hibernate가 모르는 타입 - 커스텀 UserType
또는 네이티브 쿼리 / 난이도 높음"으로 적어뒀으나, Hibernate 6.4부터
hibernate-vector 모듈이 SqlTypes.VECTOR를 공식 지원한다.
현재 7.2.0.Final이고 Spring Boot BOM이 버전까지 관리해 의존성 한 줄로 끝난다.
조사 없이 난이도를 추정한 결과였다.
@JdbcTypeCode(SqlTypes.VECTOR)
@array(length = EMBEDDING_DIM)
private float[] embedding;
byte[] <-> float[] 변환 단계가 운영 경로에서 사라졌다(ChunkIndexWriter,
RetrievalService). EmbeddingClient.toBytes/toFloats는 Phase 1 기준선을 재현하는
BruteForceSearchBenchmarkTest만 쓰므로 그 사실을 주석에 남겨 남겨뒀다.
## 예상 밖이었던 것
1. 차원이 스키마가 됐다
HNSW는 고정 차원 컬럼에만 걸린다. 그 대가로 차원이 다른 모델의 청크가
공존할 수 없다 - 가변 길이 BYTEA 시절에는 가능했던 성질이고,
embedding_dim을 둔 근거("모델 교체 과도기") 자체가 성립하지 않게 됐다.
2. 확장 등록은 이미지와 별개다
pgvector 이미지에 확장 파일은 있으나 CREATE EXTENSION은 따로 실행해야 한다.
없으면 vector(384) 컬럼 생성이 실패하는데 ddl-auto=update가 DDL 오류를
로그만 남기고 기동을 막지 않아 정상 기동한 것처럼 보인다 -
2a에서 SafetyCourse의 DATETIME이 숨었던 것과 같은 방식이다. 두 번째다.
db/init/을 마운트했으나 데이터 디렉터리가 빈 경우에만 돌므로,
기존 볼륨용으로 docs/migration/2b-pgvector.sql을 따로 뒀다.
3. ALTER가 행 0건에서도 USING을 요구한다
bytea -> vector 자동 변환 규칙이 없어 DDL 시점에 거부된다.
USING NULL::vector(384)로 값을 버리겠다고 명시했고, 이 식은 행이 남아 있으면
NOT NULL 위반으로 실패해 실수로 임베딩을 날리는 것을 막아준다.
4. 벡터가 TOAST로 나간다
vector(384)는 1,544 bytes이고 attstorage가 EXTERNAL이다.
본문과 합쳐 행이 약 2KB를 넘으면 벡터가 행 밖으로 빠진다
(실측: 비압축성 본문 1,000행에서 TOAST 힙 1,600kB, 나간 값의 길이가
1,540 bytes로 content가 아닌 벡터였다. 본문이 압축되면 인라인으로 남았다).
ADR-0006이 VARBINARY를 고르며 피하려던 InnoDB 오프페이지 저장과 같은 문제가
PostgreSQL에서 재현된 것이다. 대응은 SET STORAGE PLAIN이지만,
HNSW 도입 후에는 인덱스가 벡터 사본을 가져 영향이 줄어들므로
근거 없이 먼저 튜닝하지 않고 인덱스 단계에서 함께 판단한다.
## 검증
VectorTypeMappingTest 신규 - 엔티티 왕복 / ChunkVector 투영 왕복 / 차원 불일치 거부.
투영을 따로 보는 이유는 검색 경로가 엔티티가 아니라 생성자 투영으로 벡터만
뽑아 오기 때문이다. 거기서 되살아나지 않으면 매핑이 반쪽이다.
DB가 없을 때 skip되도록 @EnabledIfEnvironmentVariable을 클래스에 붙였다.
기존 HibernateBatchInsertVerificationTest는 Assumptions가 메서드 안에 있어
컨텍스트 로딩이 먼저 실패한다 - 별도로 정리할 대상이다.
HibernateBatchInsertVerificationTest는 vector(384)에서도 그대로 통과한다
(10,000건 / nextval 200회). RAG 테스트 21건 통과, 1건 skip(임베딩 서버 미기동).
## 알려진 잔여
schema.sql은 2a에서 이관되지 않아 여전히 MySQL 문법이다
(AUTO_INCREMENT, ENUM, SET FOREIGN_KEY_CHECKS 등). 실행되지 않는 참조 문서라
드러나지 않았다. 이번에는 document_chunks.embedding 컬럼만 맞췄고
전체 이관은 별도로 다룬다.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
document_chunks.embedding에 HNSW 인덱스를 만들고 10만 청크로 실측했다.
합성 랜덤 벡터라 지연시간은 재되 recall은 재지 않았다 - 384차원 균등 랜덤은
서로 거의 직교해 오히려 ANN에 최악 조건이다. recall은 실데이터 색인 후에 잰다.
## 가장 중요한 결과는 인덱스가 아니었다
| 경로 | 10만 청크 / top-5 |
|---|---|
| Phase 1 - 애플리케이션에서 계산 | 4,059 ms |
| 인덱스 없이 DB 안에서 정확 최근접 | 약 57 ms |
| HNSW (워밍) | 약 1.7 ms |
인덱스를 하나도 만들지 않은 상태에서 이미 약 70배다. 8절의 "병목의 81%는
DB 내부 실행이 아니라 전송"이라는 진단이 확인됐다. HNSW의 추가 이득은 그 위에
얹히는 두 번째 층이며, 계산을 옮기는 것이 인덱스를 만드는 것보다 먼저이고 더 크다.
그래서 다음 작업 순서를 바꿨다 - 인덱스 파라미터 조정이 아니라
RetrievalService를 ORDER BY <=> 로 옮기는 것이 우선이다.
## 권한 선필터를 걸면 HNSW가 무너진다 (6절이 예고한 지점)
일반 사용자 경로(공개 10,000 + 본인 20 = 후보의 약 10%)에서 top-5 요청:
| 방식 | 반환 | 지연 |
|---|---|---|
| HNSW 기본 | 1 / 5 | 1.6~2.3 ms |
| HNSW + iterative_scan | 5 / 5 | 3.9~5.4 ms |
| 정확 최근접 + 선필터 | 5 / 5 | 15.2~22.6 ms |
5개를 요청했는데 1개가 온다. EXPLAIN의 "Rows Removed by Filter: 39"가 원인이다 -
ef_search=40으로 40개를 훑었으나 39개가 권한에 걸렸다. HNSW 그래프는 전체
데이터로 짜여 있어 볼 수 있는 문서가 10%뿐인 사용자에게는 탐색 경로 대부분이 버려진다.
기밀성이 아니라 완전성이 깨진 것이다. 필터는 여전히 SQL 안에서 걸리므로
남의 문서는 나가지 않는다. 무너지는 것은 recall이다.
그리고 필터가 붙는 순간 HNSW 우위가 32배에서 약 4배로 줄어든다.
정확 최근접은 15~22ms에 recall 100%인데 HNSW는 4~5ms에 recall 미지수다.
일반 사용자 경로에서는 인덱스를 안 쓰는 선택이 합리적일 수 있으며,
그 판단은 실데이터 recall 측정 전에는 내리지 않는다.
## 빌드에서 걸린 두 가지
1. maintenance_work_mem 64MB에서 10만 건 중 28,363건에 그래프가 넘쳐
디스크 기반 빌드로 전환된다. NOTICE로 알려주므로 조용히 느려지지는 않는다.
4분 56초 -> 256MB + 병렬 끔으로 3분 09초.
2. maintenance_work_mem을 올리면 병렬 빌드가 실패한다.
ERROR: could not resize shared memory segment ... No space left on device
병렬 워커는 스레드가 아니라 별도 프로세스라 DSM이 필요하고 리눅스에서는
/dev/shm에 잡히는데, Docker가 컨테이너 메모리 한도와 무관하게 기본 64MB를 준다.
메시지가 디스크를 의심하게 하지만 실제로는 RAM 기반 tmpfs가 꽉 찬 것이다.
병렬을 끄니 같은 256MB로 성공한 것이 진단 근거다.
compose.yaml에 shm_size를 키우는 대안은 택하지 않았다. tmpfs 페이지도 컨테이너
메모리 cgroup에 잡히는데 한도가 512MB뿐이라 shared_buffers 128MB + DSM 256MB면
빌드 중 OOM 위험이 실재한다. 인덱스 빌드는 상시 경로가 아니다.
## 운영상 주의
인덱스가 Hibernate 관리 밖이다. @Index는 btree만 만들고 HNSW 문법을 모른다.
db/init/에 넣을 수도 없다 - 컨테이너 최초 기동 시점에는 Hibernate가 아직
테이블을 만들기 전이다. 결국 마이그레이션 스크립트로만 관리되며 새 환경에서
빠뜨리면 조용히 없는 상태가 된다. CREATE EXTENSION에 이어 두 번째다.
인덱스 크기가 195MB로 테이블(195MB)과 맞먹어 shared_buffers 128MB에 둘 다
상주할 수 없다. 콜드 캐시 첫 질의 133ms, 워밍 1.6ms. 콜드 상태에서는 HNSW가
정확 스캔보다 오히려 느렸다(239ms 대 194ms) - 워밍업 없는 측정은 신뢰할 수 없다.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
RetrievalService가 후보 벡터를 전부 가져와 자바에서 코사인을 돌리던 구조를 ORDER BY embedding <=> :q LIMIT n 네이티브 쿼리로 옮겼다. 8-2의 실측이 이 순서를 정했다 - 인덱스보다 계산 위치가 먼저다 (인덱스 없이 DB에서 계산만 해도 4,059ms -> 약 57ms). ## 2단계 왕복이 함께 사라졌다 6절은 메모리 때문에 "벡터만 투영 조회 -> 점수 계산 -> 선택된 top-k만 본문 재조회" 구조를 택했다. 후보 전체를 메모리에 올려야 해서 본문까지 실으면 10만 청크에서 300MB가 되어 Xmx400M에서 위험했기 때문이다. 계산이 DB로 넘어가면서 그 전제가 소멸했다. 애플리케이션에 오는 것은 상위 몇 건뿐이라 본문을 같이 실어도 수십 KB다. ChunkVector 투영 + findByChunkIdIn 재조회가 NearestChunk 한 번의 조회로 합쳐졌다. 제약이 사라지면 그 제약에 맞춰 만든 구조도 걷어내야 한다. ## 유지한 것 - 권한 선필터: 공개/소유를 각각 top-n으로 뽑아 합친다. 후보에 애초에 안 들어온다 - OR 분리: 각 분기가 단일 조건이라 실행계획이 단순하다 - 문서당 최대 2청크: 유지하되 방식이 바뀌었다(아래) 문서당 상한 때문에 over-fetch가 생겼다. topK만 가져오면 한 문서가 상위를 채웠을 때 상한 적용 후 개수가 모자라므로 topK * 4를 요청한다. 이 over-fetch는 6절이 거부한 그것이 아니다. 6절이 거부한 것은 권한을 나중에 거르려고 넉넉히 가져오는 방식이었고 문제는 "요청한 k보다 적게 반환되어 문서 존재가 유출된다"였다. 여기서 더 가져오는 이유는 다양성 상한이며 권한 필터는 여전히 SQL 안에서 걸린다. ## 새로 생긴 실패 경로 컬럼이 vector(384)로 고정되어 질의 벡터의 차원이 다르면 DB가 쿼리를 거부한다 (different vector dimensions). 그대로 흘리면 폴백 조건에 걸리지 않으므로 쿼리 전에 길이를 확인해 EMBEDDING_DIMENSION_MISMATCH로 바꿔 키워드 폴백으로 보낸다. 색인 모델을 바꾸고 재색인하지 않은 상태가 이 경로로 들어온다. min-score는 SQL이 아니라 애플리케이션에서 적용한다. 거리 조건을 WHERE에 넣으면 HNSW 탐색 뒤 필터로 적용되어 8-2에서 본 "반환 개수 부족"을 하나 더 만든다. ## 검증 VectorSearchQueryTest 신규(실제 PostgreSQL). 단위 테스트로는 SQL이 한 줄도 실행되지 않는다 - RetrievalServiceTest는 리포지토리를 목킹하기 때문이다. 실제 DB에 던져야만 드러나는 것들을 본다: - CAST(:queryVector AS vector) 바인딩. JDBC 파라미터에는 타입 정보가 없어 문자열로 넘기고 SQL에서 캐스팅한다 - LIMIT :limit 파라미터 바인딩 - 인터페이스 투영과 컬럼 별칭의 결합 - <=> 가 정말 1 - 코사인 유사도인가 (직교 1.0, 45도 0.293 확인) - 권한 선필터가 실제로 남의 청크를 후보에서 빼는가 RAG 테스트 28건 통과, 1건 skip(임베딩 서버 미기동). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CREATE EXTENSION과 HNSW 인덱스는 Hibernate가 만들 수 없고, 둘 다 빠뜨려도 에러가 나지 않는다. 확장이 없으면 테이블 생성이 실패하지만 ddl-auto=update가 DDL 오류를 로그만 남기고 기동을 막지 않아 정상 기동한 것처럼 보이고 (2a의 SafetyCourse 사고와 같은 방식), 인덱스가 없으면 전체 스캔으로 조용히 돌아가 몇 달 뒤에나 눈치챈다. docker-entrypoint-initdb.d로는 해결되지 않는다. 그것은 컨테이너가 빈 데이터 디렉터리를 초기화할 때 실행되므로 확장은 되지만 인덱스는 안 된다(그 시점에는 테이블이 아직 없다). 게다가 볼륨이 이미 있으면 아예 돌지 않는다. Flyway는 Spring Boot가 EntityManagerFactory보다 먼저 실행하므로 확장 -> 테이블 -> 인덱스 순서가 한 번에 보장된다. ## 범위: 점진 도입 document_chunks만 V1이 관리하고 나머지 테이블은 여전히 ddl-auto=update가 만든다. 그래서 V1의 모든 문장이 IF NOT EXISTS이며 기존 DB에도 안전하게 적용된다. ddl-auto를 validate로 낮추지 않았다. 나머지 20여 개 테이블이 아직 마이그레이션에 없으므로 빈 DB에서 validate를 걸면 기동 자체가 실패한다. ## 도중에 잡은 문제 둘 1. flyway-core만으로는 마이그레이션이 아예 돌지 않는다 Spring Boot 4에서 자동 구성이 기술별 모듈로 쪼개져 FlywayAutoConfiguration이 spring-boot-autoconfigure에 없다. 라이브러리는 클래스패스에 있는데 아무도 실행하지 않는 상태였고, flyway_schema_history가 안 생겨서 발견했다. spring-boot-flyway로 교체했다. 2. defer-datasource-initialization이 앱 기동을 막았다 Circular depends-on relationship between 'flyway' and 'entityManagerFactory'. 이 설정은 spring.sql.init(data.sql)을 Hibernate DDL 뒤로 미루는 것인데 sql.init.mode=never라 미룰 대상이 없어 아무 일도 하지 않던 잔재였다. Flyway가 들어오자 초기화 순서를 흔들어 순환을 만들었다. 주석 처리했다. @DataJpaTest만 깨진 게 아니라 @SpringBootTest(contextLoads)도 실패했다. 즉 실제 앱이 기동하지 않는 상태였다. ## 검증 FlywayMigrationTest 신규. 임시 데이터베이스를 만들어 빈 상태에서 적용한다 - 개발 DB에는 이미 확장도 인덱스도 있어서 아무것도 안 해도 통과하기 때문이다. 확장 등록 / vector(384) 컬럼 / 인덱스 접근 방식이 정말 hnsw인지(이름만 보면 btree여도 통과한다) / 시퀀스 증가폭 50 / 보조 인덱스 / 멱등성을 확인한다. 스프링 컨텍스트는 띄우지 않는다. @DataJpaTest는 Flyway를 자동 구성하지 않고, 억지로 넣으면 위와 같은 순환이 생긴다. 다만 그 순환 자체가 Spring Boot가 EMF를 Flyway에 의존시킨다는 증거이므로 순서 보장은 프레임워크가 담보한다. RAG 테스트 34건 통과(1건 skip - 임베딩 서버 미기동), contextLoads 통과. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
RAG/벡터 검색과 PostgreSQL 이관이 끝났는데 문서는 그 이전 상태로 남아 있었다. 특히 PROJECT-SCOPE는 "벡터 임베딩 기반 교차언어 검색은 이 프로젝트에 존재하지 않는다"고 명시하고 있어 구현과 정면충돌했다. - PROJECT-SCOPE: 전면 개정. "모델을 만들지 않았고 모델을 쓰는 시스템을 만들었다"는 경계를 표로 명시(가져다 쓴 것 / 직접 만든 것). 포폴에서 과대 서술하지 않기 위한 문서이므로 이 구분이 핵심이다. - ADR-0003: 6절 신설. 전제 두 가지(저장소 PostgreSQL, 다국어는 벡터로 해결)가 바뀌었음을 반영하고 Elasticsearch를 도입하지 않는 근거 정리. 3-4(정합성 설계)는 전면 이전으로 소멸했다. - ADR-0005: 본문 2~4절은 진단 시점 스냅샷이라 고치지 않고 5절에 갱신 표를 붙였다. 격차 진단을 현재 상태로 덮어쓰면 "무엇을 채웠는가"가 사라진다. - WORK-BACKLOG: 정합성 섹션을 완료로 전환, 신규 P0-7~10 등록. 점검 과정에서 문서가 반대 방향으로 틀린 항목이 나왔다. CI/CD를 "충족"으로 적어뒀으나 워크플로가 계속 실패해 삭제된 상태였다(e672ea9). 테스트는 1개에서 20개로 늘었는데 자동 실행되는 곳이 없다. 진단 당시보다 오히려 나빠진 유일한 항목이라 P0-7로 올렸다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
파일 주석에 "머신마다 다른 로컬 설정이라 커밋하지 않는다"고 적어뒀으나 정작 .gitignore에는 없어서 매번 status에 untracked로 떴다. 그 상태로는 누군가 실수로 add 할 수 있고, 로컬 포트 오버라이드가 공유 설정이 되면 다른 개발자 환경이 조용히 깨진다. application-local.properties 옆에 로컬 전용 설정으로 묶었다. 확장자 두 형태 모두 Docker Compose가 인식하므로 함께 넣었다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
테스트가 1개에서 20개로 늘어나는 동안 자동으로 돌아가는 곳이 없었다. 이전 워크플로는 배포 실패로 통째로 삭제된 상태였다(e672ea9). 백로그 P0-7은 "DB가 필요한 테스트는 DB_PASSWORD 유무로 skip되므로 CI에서도 안전하게 돌릴 수 있다"고 적어뒀으나, 환경변수 없이 실제로 돌려보니 3개가 실패했다. 아무도 검증한 적 없는 진술이었다. - DocKinSpringApplicationTests: 가드가 아예 없었다 - LockTimeoutVerificationTest: 가드가 메서드 안(Assumptions)에 있었다 - HibernateBatchInsertVerificationTest: 동일 뒤의 둘은 직전 Flyway 커밋(d3fdc15)이 깨뜨린 것이다. 그전에는 Hibernate가 커넥션을 늦게 잡아 메서드의 Assumptions가 먼저 실행됐는데, Flyway는 컨텍스트 기동 시점에 접속하므로 메서드에 닿기 전에 터진다. VectorTypeMappingTest가 주석으로 남겨둔 "가드는 클래스에 붙여야 한다"가 정확히 이 이야기였고, 셋 다 그 규약(클래스 레벨 @EnabledIfEnvironmentVariable)으로 맞췄다. 결과: 100개 중 78 통과 / 22 skip / 0 실패. CI는 테스트만 복구했다. 삭제된 워크플로가 실패한 원인은 테스트가 아니라 Docker Hub 푸시와 EC2 배포였고, 한 잡에 묶여 있어 배포 자격증명이 없다는 이유로 전체가 빨간불이 됐다. 배포를 다시 붙이더라도 별도 워크플로여야 한다. (그 워크플로는 -x test라 살아 있을 때도 테스트를 돌린 적이 없다.) CI에 DB_PASSWORD를 주지 않는 이유와 그래서 검증되지 않는 항목은 ci.yml 하단과 백로그에 적어뒀다. 후속으로 P0-11(테스트 스위치 분리), P0-12(컨텍스트 로딩 검증 복원)를 등록했다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
복구한 워크플로가 통과는 했으나 네 액션 모두 Node 20 기반이라 런너가 강제로 Node 24에서 돌리고 있었고, setup-java v4는 업데이트 중단 안내가 붙었다. 방치하면 런너가 Node 20을 걷어내는 시점에 CI가 다시 죽는다 - 워크플로가 삭제됐던 것과 같은 경로다. checkout v4->v7, setup-java v4->v5, upload-artifact v4->v7, setup-gradle v4->v6. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
작업일지 4,000건/일(100만 건/년), 청크 1년 200~300만인데 만료 개념이 없다.
현재 삭제되는 청크는 재색인 시 stale 정리뿐이고 그건 멱등성을 위한 것이지
보존 정책이 아니다. ADR-0006은 이 증가를 Phase 2 전환 트리거로만 다뤘고
2년차 이후를 다룬 적이 없다.
결정:
- 원본(work_logs)은 TTL을 걸지 않는다. 법정 보존기간 확인이 선결이며
확인 전까지는 되돌릴 수 없는 쪽을 택하지 않는다.
- 색인(document_chunks)만 최근 N개월로 제한한다. 파생 데이터라 재생성이
가능하고, 용량의 대부분이 여기 있다(HNSW 인덱스가 테이블만 한 사본을
하나 더 갖는다 - 8-2 실측).
- 만료는 DELETE가 아니라 range 파티션 + DROP PARTITION. DELETE는 dead
tuple과 HNSW 그래프 열화를 남긴다. 파티션은 인덱스가 통째로 사라진다.
작성하면서 스키마 결함을 하나 찾았다. DocumentChunk.createdAt은 청크가
색인된 시각이지 원본 작성일이 아니다. 모델을 교체해 전체 재색인하면 전부
리셋되어 3년 전 작업일지가 "오늘 것"이 되고 만료를 빠져나간다. 지금은
처음 색인하는 중이라 두 값이 우연히 비슷해 드러나지 않는다 - 보존 정책을
붙이는 순간 조용히 틀리는 종류다. source_created_at 비정규화가 선결 작업이
되는 이유이며, 백로그 P2-7-1로 올렸다.
만료 기간 N은 정하지 않았다. chat_history에 채택 근거가 기록되므로 그
원본 작성일 분포를 그리면 나오지만 실사용 로그가 아직 없다. 분포가
평평하면 이 문서의 전제("오래된 문서는 덜 유용하다") 자체가 틀린 것이라
그 경우도 적어뒀다.
파티셔닝의 대가(복합 PK, 파티션별 인덱스로 인한 전 기간 검색 저하)와
검증이 필요한 항목은 [검증 필요]/[측정 필요]로 표시했다.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
P0-8: 중복 선언이 아니었다. DocKinSpringApplication에는 @EnableJpaAuditing 애노테이션이 없고 사용하지 않는 import만 남아 있었다. 선언은 처음부터 JpaAuditingConfig 한 곳뿐이다. import를 제거했다. P0-9: BusinessException 하나만 처리해 @Valid 실패가 500으로 나갔다. 검증/타입/본문파싱/메서드/권한/캐치올 핸들러를 추가했다. - 응답 형태는 바꾸지 않았다. ErrorResponseDto에 ErrorCode의 코드값을 넣으면 프론트엔드 계약이 바뀌므로 별개 결정으로 남긴다. - 캐치올을 넣는 순간 AccessDeniedException이 삼켜져 403이 500이 된다. RuntimeException이기 때문이다. 명시적 핸들러로 막고 주석을 남겼다. - 검증 실패 응답에 입력값을 넣지 않는다(비밀번호가 응답·로그에 남는다). 500에도 예외 메시지를 노출하지 않는다(파서 메시지에 패키지 구조가 있다). P0-10: SENTENCE_LOOKBACK=120을 실측했다. SentenceLookbackMeasurementTest로 탐색 거리를 스윕하되, 강제 절단 비율만 재면 답이 "최대한 키워라"로 나오므로 평균 청크 길이를 함께 쟀다. 측정 대상과 다른 코드를 재지 않도록 프로덕션 경로에서 세게 splitWithStats를 추가했다. 가설이 하나 틀렸다. "창을 넓히면 청크가 짧아진다"고 보고 어서션을 걸었는데 실패했다. 탐색이 뒤에서부터 훑어 가장 가까운 경계를 쓰기 때문에 이미 경계를 찾은 절단은 창을 넓혀도 위치가 그대로다. 작업일지 문투에서 120과 400이 완전히 동일했다(448자, 401청크). 대가는 강제 절단이 실제로 구제되는 경우에만 붙는다(긴 서술 120->200에서 평균 8% 감소). 값은 유지하고 근거를 붙였다. 필요한 거리는 평균 문장 길이의 약 1.7배이며 (35자->80, 70자->120, 140자->200), 즉 120은 "평균 문장 길이 70자" 가정값이다. 남은 실측은 값이 아니라 그 가정이다. STT 텍스트에서는 종결 부호가 없어 어떤 값이든 100% 강제 절단이라는 것도 드러났다. 이 프로젝트에 STT 경로가 있으므로 해당 텍스트에서 이 파라미터는 아무 일도 하지 않는다. 테스트 102개 중 80 통과 / 22 skip / 0 실패. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"PostgreSQL 스키마를 보려면 어디를 봐야 하나"에 답이 없었다. ddl-auto=update가 스키마를 만드는 구조라 파일이 존재한 적이 없고, 엔티티 22개를 읽는 수밖에 없었다. SQL 파일은 6개나 있었지만 실행되는 것은 2개뿐이고 나머지는 MySQL 시절 잔재라 오히려 오해를 부르는 상태였다. - docs/db/postgresql-schema.sql: 실제 DB에서 뽑은 덤프(24테이블). 생성물이며 소스 오브 트루스가 아니라는 것과 갱신 방법을 헤더에 적었다. - SchemaValidationTest: ddl-auto=validate로 엔티티와 DB의 일치를 검증한다. 파일이 없으면 비교할 대상이 없어 어긋나도 아무도 모른다는 문제의 대응이다. - schema.sql에 stale 경고 헤더. 실행되지 않고, MySQL 문법이고, 내용도 낡았다는 것이 파일 안에 아무 표시도 없었다. 검증은 통과했다. 다만 drift가 없는 이유는 보호 장치가 아니라 우연이다 - DB가 2026-08-04 23:24 생성으로 PostgreSQL 이관 때 통째로 새로 만들어져 현재 엔티티와 일치할 수밖에 없었다. 앞으로 컬럼을 지우거나 타입을 바꾸면 update는 반영하지 않으므로 그때부터 어긋난다. validate의 한계도 적어뒀다. 테이블/컬럼 존재와 타입만 보고 nullability, 기본값, 인덱스, FK 제약, 그리고 DB에만 있는 잉여 항목은 잡지 못한다. 실제로 bench_leave_balance(LeaveBalanceConcurrencyTest가 만드는 벤치마크 테이블)가 잉여로 남아 있으나 검증을 통과한다. 남은 부채는 P0-13-1/2/3으로 등록했다. 근본 해결은 전체 테이블을 Flyway로 옮기고 ddl-auto=validate로 내리는 것이며, 지금은 테스트가 그것을 대신하고 있어 기동 자체는 여전히 막지 못한다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
552줄짜리 MySQL 8.0.44 덤프(Dockin 데이터베이스)가 저장소 루트에 남아 있었다. compose.yaml, Dockerfile, application.properties, gradle 어디에서도 참조하지 않는 고아 파일이다. 루트에 있고 이름이 init.sql이라 저장소를 처음 여는 사람이 스키마 정의로 오인하기 가장 쉬운 파일이었다. 실제 스키마는 PostgreSQL이고 엔티티에서 생성되며, 읽을 파일이 필요하면 docs/db/postgresql-schema.sql이다. 필요하면 b67134f에서 복구할 수 있다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
LocalMigrationDriftTest를 추가했다. 저장소의 마이그레이션이 로컬 DB에 실제로 적용되어 있는지 본다. 앱 안에 넣을 수 없는 검사다. 옛 이미지의 앱 입장에서는 jar에 V1/V2, DB에 V1/V2로 완벽히 일치했다. 어긋난 것은 저장소와 DB 사이이고 앱은 저장소를 모른다. 이 비교는 작업 트리를 읽을 수 있는 쪽에서만 성립하므로 테스트다. 새 환경변수 스위치를 만들지 않았다. 이 저장소에는 DB_PASSWORD로 켜지는 로컬 전용 테스트가 이미 여럿이고 CI 워크플로가 "그 스위치가 이름과 실제가 어긋난 상태"라고 스스로 적어 두었다. 스위치를 더하면 켜는 것을 사람이 기억해야 하는 장치가 되는데, 그건 이 테스트가 잡으려는 병과 같은 병이다. 그래서 조건 자체를 스위치로 쓴다. .env가 있는가(.gitignore에 있으므로 개발 머신에만 있다)와 localhost:5432에 붙는가다. 개발 머신에서는 둘 다 자연히 참이고 CI에서는 둘 다 자연히 거짓이라 따로 켤 것이 없다. CI의 DB는 Testcontainers가 매번 새로 만들어 정의상 밀릴 수 없으므로 거기서 건너뛰는 것이 맞다 - 항상 통과하는 검사는 아무것도 검증하지 않는다. 검사는 둘이다. 밀린 쪽(pending)과 앞선 쪽(MISSING_SUCCESS/FUTURE_SUCCESS)이고, 뒤쪽은 앱을 띄워야 알 수 있던 "Detected applied migration not resolved locally"를 띄우기 전에 말해 준다. info()만 쓰고 migrate()는 쓰지 않는다 - 테스트가 개발 DB의 스키마를 바꾸기 시작하면 그 자체가 사고의 원인이 된다. 감지하는지 확인했다. 임시 V99를 넣고 돌린 뒤 지웠다. 정상 (저장소 V5 = DB v5) PASSED 밀림 (V99 추가) FAILED - 밀린 버전과 고치는 명령이 함께 나온다 .env 없음 (CI 조건) SKIPPED, BUILD SUCCESSFUL 세 번째를 확인한 이유는 여기서 skip이 아니라 fail이 나면 CI를 깨뜨리기 때문이다. flyway_schema_history는 실험 전후로 0,1,2,3,4,5 그대로다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
루트의 app.jar 87,630,035바이트를 지우고, 다시 들어오지 못하게 .gitignore에 *.jar를 넣고, 도커 빌드 컨텍스트를 .dockerignore로 잘랐다. a22eacb에서 들어간 뒤 아무도 읽지 않은 파일이다. .gitignore가 막고 있던 것은 build/ 하나뿐이라 루트로 떨어진 jar는 애초에 규칙 밖이었다. 넣은 쪽이 잘못한 것이 아니라 넣으면 걸러 주는 것이 없었다. 지워도 되는지부터 봤다. Dockerfile은 build/libs에서 COPY하고, 저장소에서 app.jar를 가리키는 자리는 이미지 안 경로 /app.jar 둘(COPY 대상과 ENTRYPOINT의 -jar)뿐이다. 루트 파일을 쓰는 곳은 없다. *.jar는 예외 줄보다 위에 뒀다. gitignore는 마지막 규칙이 이기므로 순서가 뒤집히면 !gradle/wrapper/gradle-wrapper.jar가 같이 죽는다. git check-ignore -v gradle/wrapper/gradle-wrapper.jar가 exit 1이다 - 안 걸린다. .dockerignore는 컨텍스트가 docker build마다 통째로 데몬에 가기 때문이다. 없을 때 292MB였고 Dockerfile이 실제로 쓰는 것은 build/libs의 jar 하나다. 다만 build를 뺀 뒤 !build/libs로 되살리는 것이 실제로 도는지는 이번에 확인하지 못했다 - 도커 데몬이 떠 있지 않았다. 틀렸다면 다음 빌드가 COPY failed로 바로 죽으므로 조용히 새지는 않는다. .git은 이걸로 안 줄어든다. 86MB의 대부분이 이 blob인데 삭제는 HEAD에서만 뺀 것이고 a22eacb의 blob은 그대로다. 히스토리 재작성은 push 안 된 커밋 8개와 다른 브랜치가 걸려 있어 따로 판단할 일이라 백로그에 남겼다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
직전 커밋에서 확인하지 못한 채 남겼던 것이다 - 도커 데몬이 떠 있지 않았다. 띄우고 돌렸다. 확인 대상은 build를 뺀 뒤 !build/libs로 되살리는 것이 도는가였다. 부모를 제외한 뒤 자식을 되살리는 것은 gitignore에서 안 되는 일이라 도커에서도 안 되면 COPY가 죽는다. 돈다. #4 transferring context: 115.47MB 55.4s done #7 [3/3] COPY build/libs/...jar app.jar DONE 2.3s 115.47MB는 사실상 jar 하나다 - build/libs의 jar가 115,441,942바이트다. 컨텍스트에 그것 말고 실린 것이 없다. 세 숫자가 맞는다. 디스크의 저장소 전체가 지금 208MB이고 지운 app.jar 84MB를 더하면 292MB - .dockerignore 전에 재 뒀던 그 숫자다. "없을 때"를 도커로 다시 재려던 쪽은 실패해서 그대로 적었다. .dockerignore를 잠시 치우고 재니 93.98MB가 나왔는데 이건 위 빌드보다 작다. BuildKit이 컨텍스트를 빌드 사이에 캐시해서 직전에 보낸 jar를 다시 세지 않았고 처음 보는 .git 86MB만 실린 것이다. --no-cache는 레이어 캐시에만 걸리고 컨텍스트 전송에는 안 걸린다. 그래서 표의 그 줄은 도커 숫자가 아니라 du 숫자로 뒀다. 빌드는 dockin-ctx-check:tmp로 따로 태그해 khyojae/dockin-app:latest를 건드리지 않았고 확인 뒤 지웠다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
직전 커밋은 실측 절을 새로 붙이기만 했고, 이미 숫자가 적혀 있던 자리들은 옛 값을 그대로 들고 있었다. .dockerignore 주석: "없을 때 292MB" 옆에 넣은 뒤 115.47MB를 적었다. jar 자체가 115,441,942바이트라 그것 말고 실린 것이 없다는 뜻이 숫자에서 바로 읽힌다. .dockerignore의 app.jar 줄: "굴러다니는"을 "굴러다니던"으로 고쳤다. 지운 뒤라 사실이 아니게 됐다. 그런데도 이 줄을 남기는 이유를 적었다 - .gitignore의 *.jar가 커밋되는 것은 막지만 빌드 전에 손으로 떨어뜨린 파일은 작업 트리에 남고, 그건 컨텍스트에 실린다. 백로그 P2-15-9의 표: "없을 때 292MB"를 "292MB -> 115.47MB"로 바꿨다. 전후가 같은 칸에 있어야 이 줄이 무엇을 했는지가 표에서 끝난다. 백로그 P2-15-5 설계 절: "뒷정리에서 걸려 숫자를 아직 못 옮겼다"가 남아 있었다. 숫자는 8일에 같은 문서 위쪽 「결과」 표로 옮겨졌다. 옮겼는데 못 옮겼다고 적혀 있으면 다음에 이 문서를 여는 사람이 없는 일을 찾는다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
8일 실행의 100만 건 "인덱스 있음" 열은 신뢰할 수 없었다. ③④⑤가 0.6~0.7배로, 인덱스를 만든 뒤가 더 느렸다. 조건을 맞춰 다시 쟀다. 없음 -> 있음을 연달아 재는 구조가 원인이었다. 먼저 돈 "없음"이 캐시를 데워 놓고 "있음"이 그 위에서 돌았다. WARMUP_RUNS로는 못 고친다 - 그건 그 쿼리가 실제로 건드린 블록만 데우므로, 테이블 전체를 읽는 ②와 20행만 보는 ①이 서로 다른 상태에서 출발하는 것은 그대로다. 쿼리와 무관하게 대상 전체를 올려야 한다. 각 패스 앞에 pg_prewarm(rel, 'read')을 넣었다. 힙 둘과 거기 달린 인덱스 전부를 올린다 - 인덱스를 빠뜨리면 "있음" 쪽만 찬 상태로 재게 되어 편향이 방향만 바뀐다. 'buffer'가 아닌 이유. shared_buffers 128MB인데 100만 행 work_logs는 그보다 크다. 'buffer'로 올리면 뒤쪽이 앞쪽을 밀어내며 들어가 끝난 뒤 남는 것은 테이블 꼬리뿐이다. 다만 'read'도 전부 올리지는 못한다 - prewarm이 약 870MB인데 DB 컨테이너 한도가 512MB고 cgroup v2에서 페이지 캐시는 그 한도에 함께 계상된다. 그래서 이 표는 "전부 캐시된 상태"가 아니라 "양쪽 패스가 같은 조건에서 출발한 상태"다. 노린 것은 대칭이다. 100만 결과 (ms, 5회 평균, 복합 없음 -> 있음) ① 1페이지 3.6 -> 2.1 1.8배 ② COUNT 85.3 -> 41.5 2.1배 ③ 500페이지 150.1 -> 100.1 1.5배 ④ 1페이지 정렬 2781.0 -> 2265.1 1.2배 ⑤ LIKE %kw% 432.4 -> 571.4 0.8배 ⑥ 이미지 ×20 23.4 -> 22.6 1.0배 ⑥이 이 표를 믿을 수 있게 하는 줄이다. P2-15-4로 work_log_images(work_log_id)가 스키마에 들어간 뒤로 ⑥의 두 열은 같은 상태를 두 번 재는 것이라 1.0이 나와야 하고 1.0이 나왔다. 8일의 역전도 사라졌다. 같은 실행의 10만 구간은 버렸다. ⑥이 0.5배(35.7 -> 75.4)다. 같은 것을 두 번 쟀는데 두 번째가 2배 느리면 두 번째 패스 전체에 벌점이 붙은 것이고 나머지 다섯 줄도 못 읽는다. 그 구간은 도커 데스크톱을 막 띄운 직후 다른 컨테이너가 올라오는 중에 돌았다. ⑥을 보지 않았다면 열 줄 중 다섯 줄을 결과로 적었을 것이다. P2-15-8이 닫혔다. ④의 실행계획을 "있음"에서 뽑으니 후보 인덱스가 있는데도 플래너가 안 골랐다. 같은 구역 84명이 IN에 들어가고 그게 100만 중 164,638행, 약 16%다. 84개 구간을 훑어 모으는 것보다 병렬 순차 스캔 + top-N 힙정렬이 싸다고 본 것이고 Buffers: read=84183이 말하듯 어차피 테이블 대부분을 읽는다. 1.2배는 인덱스가 쓰여서 난 차이가 아니다 - 계획에 인덱스가 없다. 보류의 이유가 "모른다"에서 "안 쓰인다"로 바뀌었고 결론은 같다: 넣지 않는다. 측정 환경을 표 위에 박아 뒀다. AWS에서 다시 재지 않는 근거도 같은 자리에 적었다 - 이 표가 주장하는 것은 절대 지연이 아니라 배수와 실행계획이고, 인덱스를 쓸지는 플래너가 선택도를 보고 정하는 문제라 호스트를 바꿔도 방향이 안 바뀐다. 벤치 흔적은 남지 않았다. 실행 전후로 work_logs 20,016행 / users 504명 그대로다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
dev의 CI가 2026-08-07부터 닷새 동안 실패하고 있었다. 오늘 push한 커밋 12개의 CI를
확인하다 드러났다 - 이번 push가 만든 회귀가 아니다. 8월 7일 실행(31188865740)과
오늘 실행(31584824093)의 실패 지점과 스택트레이스가 같다.
DocKinSpringApplicationTests > Failed to load ApplicationContext
Caused by: RedisConnectionException: Unable to connect to Redis server: localhost:6379
P2-11-5가 연 문이다. DB를 Testcontainers로 옮기기 전까지 이 테스트는 DB_PASSWORD
부재로 스스로 skip했다. 컨테이너가 생기면서 CI에서 처음 진짜로 실행됐고 전체
컨텍스트가 필요해졌는데, 그 안의 RedissonConfig는 빈 생성 시점에 localhost:6379로
실제로 접속한다. DB만 주고 Redis는 안 줬다. application.properties의 기본값이
${REDIS_HOST:localhost}라 "설정이 없다"가 아니라 "런너의 6379에 붙으려 든다"가 되고,
거기 아무것도 없으니 죽는다.
Redis도 Testcontainers로 띄운다. 후보 셋 중에 골랐다.
services: 컨테이너 워크플로 몇 줄로 가볍지만 DB는 코드가, Redis는 워크플로가
띄우게 되어 같은 종류가 두 곳에 갈린다
RedissonClient 목 제일 빠르지만 이 테스트의 목적이 "컨텍스트가 실제로 뜨는가"다.
목을 끼우면 뜨지 않아도 통과하므로 잡으려던 것을 못 잡는다
Testcontainers DB와 같은 방식이라 "인프라는 컨테이너가 준다" 하나로 설명된다.
운영과 같은 redis:7-alpine이라 분산락이 실물에서 검증된다
PostgresTestSupport를 ContainerTestSupport로 바꿨다. Redis를 띄우는 클래스가
Postgres...라는 이름을 달고 있으면 이 저장소가 여러 번 지적한 "이름과 실제가 어긋난
상태"가 된다. 참조 10개는 기계적 치환이다.
Redis에는 설정을 주지 않았다. compose.yaml의 Redis도 기본 설정으로 뜨고, 여기서
기대하는 것이 분산락 SETNX 수준이라 튜닝할 파라미터가 없다. DB에 lock_timeout을 준
것과 대비되는데 그건 없으면 테스트가 실패가 아니라 멈추기 때문이었다.
확인. 이 머신에 실행 중인 Redis 컨테이너가 0개다(dockin-redis는 3일 전 종료).
localhost:6379에 아무것도 없는데 컨텍스트가 떴으므로 Testcontainers 쪽에 붙은 것이다.
전체 스위트 34개 클래스 failures=0 errors=0, skip은 임베딩 서버와 벤치 스위치뿐이다.
분산락 테스트(LeaveBalanceConcurrencyTest)도 2/2 통과한다.
백로그에 P2-16으로 적었다. 남은 것은 이번 건의 진짜 교훈이다 - 닷새 동안 아무도
몰랐고 다른 일을 하다 우연히 봤다. 빨간불이 조용한 상태를 없애지 않으면 같은 일이
반복된다. P0-12의 "컨텍스트 검증이 CI에서 돈다"도 그때까지는 "돌지만 아무도 안 본다"가
맞다.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
P2-16이 남긴 것이다. 닷새가 걸린 이유는 원인의 난이도가 아니라 아무도 보지
않았다는 것이었다. 다른 일을 하다 우연히 봤다.
메일은 이미 있었다. GitHub는 실패할 때마다 보내고 있었고 그것으로 부족하다는
것이 그 닷새로 증명됐다. 그래서 붙인 것은 알림 채널이 아니라 남는 형태다 -
메일은 읽지 않으면 사라지지만 열린 이슈는 저장소를 열 때마다 보인다.
푸시가 실패하면 이슈가 열리고 다시 통과하면 닫힌다.
푸시에만 건다 PR의 실패는 PR 화면에 이미 빨갛게 보인다. 이슈까지
만들면 같은 것을 두 번 말하고, 두 번 말하면 안 보게 된다
이슈 하나를 재사용 연속으로 깨질 때 커밋 수만큼 이슈가 생기면 알림이 다시
소음이 된다. 열린 것이 있으면 댓글만 단다
푸시한 사람에게 배정 팀 저장소에서 "누군가 봐야 한다"는 아무도 안 본다는
뜻이 되기 쉽다. 배정은 메일과 달리 받는 사람이 정해진다
닫는 것까지 한 쌍 고친 뒤에도 열려 있으면 다음 빨간불이 "이미 열려 있네"에
묻힌다. 여닫이가 맞아야 이 신호를 계속 믿을 수 있다
취소는 알리지 않음 concurrency가 이전 실행을 취소하는데 그때마다 이슈가
열리면 연속 푸시가 곧 소음이 된다
permissions: issues: write를 명시했다. 이 저장소의 기본 토큰 권한이 read다
(actions/permissions/workflow로 확인). 없으면 잡은 정상으로 보이는데 이슈만
안 생긴다 - 이 저장소가 여러 번 잡아온 "설정은 있는데 안 먹는다"와 같은 자리다.
아직 실패 경로를 실제로 돌려보지 않았다. 이 커밋은 dev에 올라가도 성공 경로만
지나가므로, 이슈가 정말 열리는지는 일부러 깨뜨려 봐야 안다. 그때까지 이 항목은
"붙였다"이고 "된다"가 아니다.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
직전 커밋은 "붙였다"까지였고 실패 경로를 돌려본 적이 없다고 그 메시지에 적어 뒀다. 돌려봤다. 임시 브랜치를 트리거에 넣고 exit 1 한 줄로 깨뜨렸다가 다음 커밋에서 그 줄만 빼 되돌렸다. 여는 것과 닫는 것을 한 번씩 실제로 지나가게 한 것이다. 실패 -> 이슈 열림 #34 생성. 라벨 🚨 CI, 담당자 Khyojae, 본문에 커밋 해시와 실행 URL 복구 -> 이슈 닫힘 #34 CLOSED + "다시 통과했다" 댓글 스텝 게이팅 실패 실행에선 복구가, 성공 실행에선 라벨 준비와 실패->이슈가 skipped 토큰 권한 기본이 read인데 issues: write 명시로 실제 생성됐다. 우려한 자리가 아니었음이 확인됐다 닫는 쪽까지 본 이유는 여는 것만 확인하면 반쪽이기 때문이다. 닫히지 않으면 다음 빨간불이 "이미 열려 있네"에 묻히고, 그때는 알림이 있다고 믿고 있어서 오히려 더 오래 못 본다. 지금 고치려는 것이 바로 그 상태다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
이슈를 라벨로만 찾고 있었다. dev와 main이 동시에 깨지면 뒤엣것이 앞의 이슈에 댓글만 달고, 한쪽이 복구되면 다른 브랜치의 이슈까지 닫힌다. 검증은 브랜치 하나로 했기 때문에 이 경우를 지나가지 않았다. 제목을 키로 쓴다. 라벨을 브랜치 수만큼 만들면 라벨 목록이 브랜치 목록이 되고, 지운 브랜치의 라벨이 남는다. 여는 쪽과 닫는 쪽이 같은 키를 봐야 한 쌍이 성립하므로 양쪽 다 TITLE로 맞췄다. --search를 쓰지 않았다. 그쪽은 검색 색인을 거쳐서 방금 만든 이슈가 아직 안 잡힐 수 있다. 목록을 받아 제목이 정확히 일치하는 것만 고르면 색인과 무관하다. 라벨 준비 스텝의 조건을 뗐다. 성공 경로도 --label을 걸어 조회하므로 라벨이 없으면 거기서도 실패한다. 실패했을 때만 만들면 라벨이 생기기 전의 첫 성공에서 깨진다. 검증용 브랜치 트리거는 되돌렸다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
dev가 빨간불이 됐다. LeaveBalanceConcurrencyTest의 음성 대조군이 두 번 연속 expected: <2> but was: <1>로 실패했다. 31587328167 (51466ff) PASSED 31589820219 (725405c) FAILED 31590095891 (8d9d296) FAILED 그 사이 자바 코드 변경은 0이다. 세 커밋 전부 워크플로와 문서다. 즉 코드가 아니라 타이밍이다. start 래치는 출발만 맞춘다. 겹치기는 기대할 뿐이고 주석도 "겹치는 구간을 최대화한다"고만 적혀 있었다. 2코어 러너가 바쁘면 한 스레드가 읽기·쓰기·커밋까지 끝낸 뒤에 다른 스레드가 스케줄되고, 그쪽은 이미 줄어든 잔액을 읽어 정상적으로 거절한다. approved가 1이 된다. 이게 실패보다 나쁜 쪽으로도 간다 - 통과할 때조차 무엇을 보인 것인지 알 수 없다. "락이 없으면 깨진다"를 보이려는 테스트가 깨지는 조건을 만들지 못한 채 통과할 수 있으면, 그건 항상 통과하는 검사와 같은 자리에 있다. 읽기 뒤에 배리어를 하나 더 둔다. 둘 다 옛 값을 읽은 것이 확인된 뒤에만 쓰기로 넘어가므로 lost update가 구조로 재현된다. 배리어는 락 없는 쪽에만 건다. 락을 쓰면 두 번째 스레드가 SELECT ... FOR UPDATE에서 막힌 채 첫 번째의 커밋을 기다리는데, 거기에 배리어를 걸면 먼저 읽은 쪽은 배리어에서 뒤엣것은 락에서 서로를 기다려 교착이 된다. 락 있는 쪽은 막히는 것 자체가 확정적이라 배리어가 필요하지도 않다. 배리어 실패는 따로 들고 나간다. 기존 catch가 전부 삼키고 있어서 "겹침이 성립하지 않았다"와 "우연히 순차 실행됐다"가 같은 메시지로 보였다. 이번에 원인을 좁히기 어려웠던 이유가 그것이다. 도커를 내린 뒤라 로컬 실행은 못 했다. 컴파일만 확인했고 검증은 CI에서 한다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
로컬(i3-6100 2코어 / 컨테이너 512MB)에서 막혀 있던 측정을 밤 단위로 편성했다. 실험 아홉 개를 세 갈래로 나눈다. A 제약 해제형 2코어·512MB가 실험 자체를 성립 못 하게 하는 것 (E1~E3) B 호스트 독립형 결과가 recall·비율·개수라 옮겨도 안전한 것 (E4~E9) C 올리지 않음 옮기면 측정 대상이 사라지거나 이미 안 하기로 한 것 C가 이 문서의 절반이다. P2-15-5가 목록 API 벤치를 AWS에서 다시 재지 않기로 한 근거를 그대로 쓴다 - 주장하는 것이 절대 지연이 아니라 배수와 실행계획이면 호스트를 바꾸는 것은 캐시 오염을 잡으려다 네트워크와 스토리지 변수를 들이는 일이다. 힙 밖/OOM 재현도 뺐다. 재려는 것이 512MB 제약 그 자체라 인스턴스를 키우면 측정 대상이 없어진다. 계획 전체를 지배하는 것은 색인 처리량이다. 6-11 실측(2.34청크/s)으로 20만 청크가 23.7시간이라 하룻밤에 안 들어간다. A1 이후 규모 실험이 멈춰 있던 이유가 "재고 싶은데 안 쟀다"가 아니라 코퍼스가 안 만들어졌다는 것이었다. 그래서 8 vCPU(m7i.2xlarge)로 간다. 밤 하나에 20만을 넣으려면 약 3배가 필요하고 TEI 상한 2.0 → 6.0이 정확히 코어 3배다. 즉 이 편성은 "TEI가 선형 스케일한다"에 건 것이고 E1이 첫 밤에 정산한다. 어긋나 봐야 4 vCPU 대비 1만원이다. 6.0은 호스트를 넘긴다. TEI 외 상한 합이 3.0이라 경계는 5.0이고, 6.0에서는 스케줄러가 실효 상한을 정한다. 빼지 않고 남기되 그렇게 읽어야 한다는 것을 미리 적었다 - 6-8이 관찰은 맞고 해석이 틀렸던 자리가 여기다. 초안에 두 개가 빠져 있어 밤 0을 새로 넣었다. 빈 코퍼스에서 시작한다고 적었는데 비어야 하는 것은 document_chunks이고 work_logs는 미리 차 있어야 한다. 그리고 밤 1의 전제인 "번역본 off"를 끌 수단이 없었다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
indexAll()이 indexTranslations()를 무조건 불렀다. 그래서 A2("번역본 색인의 이득을
잰다")를 재려 해도 끄는 쪽을 만들 수 없었다.
번역본 색인은 교차언어 검색의 성립 조건이 아니다. CrossLingualRetrievalTest가
한국어 코퍼스만 두고 영어 5/5, 베트남어 4/5로 통과한다 - 다국어 모델을 쓰는 이상
번역본이 없어도 교차언어는 성립한다. 성립 조건이 아니라 품질 보강이고, 보강에는
값이 붙는다. 저장·임베딩이 배로 늘고 원문과 번역본이 top-k 자리를 중복으로
차지한다(P2-8-2에 실측 재현이 있다 - 영어 질의 top-5에 작업일지 7이 두 번 들어왔다).
이 스위치를 측정용 임시 장치로 보지 않는다. A2가 "이득 작음"으로 나오면 내릴 조치가
정확히 번역본 색인을 끄는 것이고, 그때 쓸 수단이 이것이다. 결론의 실행 수단을 미리
만드는 셈이라 측정이 끝나도 남는다. 그래서 rag.indexing.enabled의 형제로 두고
기본값을 true로 뒀다 - 이득을 재기 전까지는 지금 동작이 맞다.
끌 때 조용히 건너뛰지 않고 로그를 남긴다. 검색 결과가 달라지는 설정이라 나중에
"왜 번역본이 안 잡히지"가 되면 그 줄이 답이 된다. chat_history.retrieval_mode에
VECTOR/KEYWORD/NONE을 남겨 "요즘 답변이 이상하다"를 "언제부터 폴백이었다"로 바꾼
것과 같은 판단이다.
부수적으로 밤 3의 성격이 바뀐다. 번역본 행은 DB에 있고 색인만 꺼져 있으므로,
켜고 다시 돌리면 해시 멱등으로 작업일지를 전부 건너뛰고 번역본만 색인한다.
전체 재색인이 아니라 약 30% 증분이라 밤 하나가 줄어든다.
검증은 컴파일까지다. off 경로는 밤 1에서 처음 돈다.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
E1은 TEI 상한 2.0~6.0을 1.0 간격으로 훑고, E8은 그 색인이 도는 동안 옆에서 구간별 처리량을 기록한다. E8이 "적재에 계측만 얹으면 공짜"였는데 얹을 계측이 없었다 - 없으면 밤 1이 E1만 내놓는다. e1-tei-cpu-sweep.sh 컨테이너를 다시 만들지 않고 docker update --cpus로 quota만 바꾼다. compose 오버레이로 하면 TEI가 재시작되어 모델 적재가 조건마다 끼어든다. 그 성공 여부를 믿지 않고 cgroup의 cpu.max를 직접 읽어 대조한다. 어긋나면 그 자리에서 멈춘다. 이 저장소가 여러 번 잡아온 "설정은 있는데 안 먹는다"가 여기서도 가능한 자리다. 조건마다 위치 P에서 5분 재고 그 슬라이스만 지워 P로 되돌린다. 코퍼스를 고정하는 것이 일차 목적이지만 두 번째 이득이 더 크다 - 모든 조건이 같은 문서를 다시 색인한다. TEI 시간은 청크 길이에 크게 좌우되므로(149자 46ms vs 317자 153ms) 조건마다 다른 문서를 색인하면 내용 차이가 상한 효과로 둔갑한다. CONDITIONS 끝의 2.0은 오타가 아니라 통제군이다. 첫 2.0과 다르면 스윕이 드리프트한 것이고 가운데 줄들의 배수를 읽을 수 없다. P2-15-5 재측정에서 통제군 6이 우연히 있었던 덕에 열 줄 중 다섯을 버릴 수 있었는데, 이번에는 의도적으로 넣었다. 색인 트리거는 새로 만들지 않았다. compose.gc.yaml이 seed 프로파일과 rag.indexing.on-startup=true를 켜므로 앱을 올리는 것이 곧 색인 시작이고, 6-11도 그 조합으로 쟀으므로 절대값 비교 가능성이 따라온다. e8-index-sampler.sh 표본마다 TEI 상한과 앱 상태를 함께 적는다. 밤 1은 색인 도중에 E1이 끼어들어 상한을 바꾸고 앱을 재시작하는데, 그러면 E8 곡선이 그 지점에서 꺾인다. 그것은 코퍼스 크기 때문이 아니라 조건이 바뀐 것이다. "E8은 P를 경계로 두 구간으로 나뉜다"가 문서에만 있으면 자고 일어나서 어디가 경계였는지 못 찾는다. 멈춰 있던 구간도 버리지 않고 표시해서 남긴다. 지우면 왜 그 시각이 비었는지 알 수 없다. *.sh text eol=lf gradlew에만 있던 규칙을 넓혔다. 지금은 이 머신의 core.autocrlf=true가 우연히 막아주고 있을 뿐이고, 그 설정이 없는 머신에서 커밋하면 CRLF가 그대로 들어간다. 그러면 EC2에서 bad interpreter로 죽는데 원인이 파일 내용에 안 보인다. 검증 한계. 문법 검사와 파서·CSV 열 수(16열)·요약 표 렌더링만 로컬에서 확인했다. docker와 DB가 필요한 경로는 한 번도 안 돌았다. 그리고 Docker Desktop에서는 컨테이너 cgroup이 호스트에 안 보여 스로틀 통제군을 볼 수 없다 - ALLOW_NO_CGROUP=1로 흐름만 확인할 수 있게 두되, 기본값에서는 통제군이 사라지므로 본 측정을 중단시킨다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
부트스트랩을 색인 시작 전에 멈춘다. 순서가 측정의 일부이기 때문이다 - 샘플러가 색인보다 먼저 떠야 E8 곡선의 왼쪽 끝이 남고, 앱은 코퍼스를 넣은 뒤에 띄워야 한다. 그래서 준비까지만 하고 다음에 칠 명령을 찍는다. 시크릿이 오늘 밤을 막지 않는다 application-secret.properties가 gitignore돼 있고 저장소에 없다. 여기서 막히면 새벽에 막힐 자리라 코드가 요구하는 값을 전수 조사했다. JWT_SECRET, JWT_EXPIRATION, JWT_REFRESH_EXPIRATION, AI_SERVER_URL, S3_* 가 기본값 없이 필요하다 - 없으면 앱이 아예 뜨지 않는다. 그런데 이 측정에 진짜 값이 필요한 것은 DB_PASSWORD 하나다. 나머지는 앱이 뜨는 데만 필요하고 색인 경로가 쓰지 않는다. 난수와 자리표시자로 만들어 쓰고 인스턴스와 함께 버린다. P2-11-6(배포 환경에 시크릿을 어떻게 넘길 것인가)은 여전히 미결이지만 밤 1의 선행 과제는 아니게 됐다. 순서에서 이유가 있는 곳 bootJar를 compose build보다 먼저 돌린다. 거꾸로면 Dockerfile이 옛 jar를 COPY하고, 그때 에러는 나지 않는다. P2-15-9가 겪은 "이미지가 옛것이면 Flyway는 정상 동작하면서 옛 상태를 유지한다"가 그 자리다. 앱을 한 번 띄웠다 내린다. Flyway가 스키마를 만들어야 코퍼스 생성기가 넣을 테이블이 생긴다. 그 시점에는 코퍼스가 비어 있어 색인이 즉시 끝난다. RAG_INDEXING_TRANSLATIONS_ENABLED=false를 .env에 박는다. 없으면 번역본까지 색인되어 E4의 기준선이 오염된다. 코퍼스 생성 뒤 청크 수를 찍는다. 0이어야 한다 - 0이 아니면 E8이 보려는 왼쪽 끝이 이미 없는 것이고, 그것을 아침에 알면 밤 하나를 버린다. usermod로 도커 그룹에 넣어도 현재 셸에는 반영되지 않는다. 재로그인을 요구하는 대신 sg docker로 자신을 다시 실행한다. 플래그로 한 번만 돌게 막았다. 검증 한계. 문법 검사만 했고 끝까지 돌려본 적이 없다. AL2023 전용이며(dnf, compose 플러그인 수동 설치) 다른 배포판에서는 1단계부터 다르다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
부트스트랩이 만드는 .env에 DB_USERNAME이 빠져 있었다.
compose.yaml은 POSTGRES_USER=${DB_USERNAME}로 DB 슈퍼유저를 만든다. 변수가 없으면
compose가 빈 문자열로 치환하고 경고만 내며, postgres 이미지는 빈 값을 미설정으로
보고 기본값 'postgres'로 만든다. 반면 앱은 compose.yaml에 리터럴로 박힌
DB_USERNAME=root로 접속한다. 즉 DB에는 root가 없고 앱은 root로 붙으려 든다.
측정 스크립트도 같이 막힌다. e1-tei-cpu-sweep.sh와 e8-index-sampler.sh가
psql -U root로 코퍼스를 세므로 표본이 통째로 db-unreachable이 된다.
틀린 곳과 터지는 곳이 멀다는 점이 나쁘다. 결번은 .env에 있는데 증상은 앱 기동
실패로 나오고, 그 사이에 compose 경고 한 줄이 있을 뿐이다. 로컬에서 드러나지 않은
이유도 같다 - 손으로 만든 .env에는 그 줄이 있다.
그래서 값을 채우는 것에서 끝내지 않고 대조를 붙였다. compose 파일에서 ${VAR}를
전부 뽑아 .env에 있는지 확인하고, 하나라도 없으면 그 자리에서 멈춘다. 사람이 두
파일을 눈으로 맞추는 대신 기계가 맞춘다. 지금 요구되는 것은 일곱 개다.
대조 로직은 로컬에서 시험했다. 완전한 .env에서 결번 없음, DB_USERNAME을 뺀 .env에서
그것만 잡힌다.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docker compose ps -q 는 실행 중인 것만 돌려준다. 중지된 컨테이너는 -qa 여야 잡힌다. 로컬에서 확인했다 - exited 상태의 dockin-redis가 -q 에서는 빈 줄, -qa 에서는 ID다. 그래서 두 곳이 깨져 있었다. E1은 조건마다 앱을 내렸다 올린다. 즉 "멈춰 있지만 존재하는" 앱의 ID가 계속 필요한데 -q 로 물었다. 게다가 부트스트랩이 앱을 내려둔 채 끝나므로, 연습 주행은 사전 점검 첫 줄에서 "서비스가 떠 있지 않다"로 죽는다. 컨테이너는 멀쩡히 있는데 그렇게 말한다. E8의 app_running도 같다. E1이 앱을 내려둔 구간에서 컨테이너를 못 찾아, "앱이 멈춤"과 "컨테이너가 없음"이 같은 답이 된다. 두 상태가 구분되지 않으면 표본의 결측 이유를 나중에 못 가른다 - 이 샘플러가 결측을 지우지 않고 표시해서 남기기로 한 이유와 어긋난다. 사전 점검도 고쳤다. 존재 여부와 실행 여부를 나눈다. 앱은 멈춰 있어도 되고(스크립트가 올린다) DB와 임베딩 서버는 떠 있어야 한다 - 없으면 창을 열어놓고 아무것도 안 늘어나는 것을 "임베딩이 느리다"로 읽게 된다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
.git이 없고 compose.yaml이 있으면 "이미 올라와 있다"로 보고 클론을 건너뛴다. 편의가 아니라 자격증명 때문이다. 클론하려면 PAT를 인스턴스에 넘겨야 하고, 그 토큰은 .git/config에 평문으로 남는다. 측정용 임시 머신에 장기 자격증명을 올리지 않는 편이 낫고, 그러면 git archive | tar 가 기본 경로가 된다. 이 방식은 추적 중인 파일만 나간다는 이득도 있다. 로컬 .env(진짜 시크릿)와 build/, .git 86MB가 따라가지 않는다. 인스턴스의 .env는 부트스트랩이 난수로 새로 만든다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
인덱스의 모드가 100644이었다. 윈도우에서 만들어진 뒤로 exec 비트가 없었고, 리눅스 체크아웃에도 그대로 없다. 측정 인스턴스에 올려 부트스트랩을 돌리자 빌드 단계에서 즉시 죽었다. scripts/aws-bootstrap.sh: line 139: ./gradlew: Permission denied 지금까지 안 드러난 이유는 아무도 리눅스에서 이 래퍼를 직접 실행한 적이 없어서다. CI는 gradle 액션이 자기 경로로 실행하고, 로컬은 윈도우라 exec 비트를 안 본다. .gitattributes가 /gradlew의 줄바꿈은 이미 고정하고 있었는데 모드는 보지 않았다 -- 같은 파일에 대해 한 축만 지키고 있었던 셈이다. 모드를 100755로 바꾸고, 부트스트랩에서도 실행권한을 확인해 세운다. 전달 경로가 tar나 zip이면 모드를 다시 잃을 수 있어 양쪽에 둔다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
인스턴스에서 실제로 띄워 보고 나온 것이다. 색인이 도는 동안 dockin-app-1의 헬스
상태가 unhealthy다. 앱은 멀쩡히 청크를 만들고 있는데(1,922개 확인) 컨테이너는
그렇게 보고한다.
/actuator/health가 db·redis·diskSpace를 모두 확인하는데, 앱 컨테이너 상한이
1.0 cpus인 상태에서 청킹과 임베딩 호출이 그것을 채우면 헬스체크(timeout 5s)가
시간 안에 돌아오지 못한다. P2-11-2가 "떴는가를 기계가 판정한다"고 붙인 그 검사가,
바쁜 것과 죽은 것을 구분하지 못한다.
스윕은 wait_healthy로 healthy를 기다리고 있었다. 그러면 조건마다 HEALTH_WAIT_MAX
(300초)를 서 있다가 죽는다. 여섯 조건이면 30분을 기다리다 아무 결과 없이 끝난다.
wait_running으로 바꿨다. 이 스크립트가 실제로 기다려야 하는 것은 "건강한가"가 아니라
"청크가 늘기 시작했는가"이고, 그 판정은 임베딩 대기 구간이 이미 하고 있다. 여기서는
프로세스가 살아 있는지만 보고, exited/dead면 로그 위치와 함께 즉시 멈춘다.
곁가지로 하나 더 드러났다. nginx가 depends_on: dockin-app: service_healthy라서
색인 중에는 뜨지 못한다("dependency failed to start: container dockin-app-1 is
unhealthy"). 부트스트랩 뒤 nginx를 올릴 때 --no-deps가 필요했다. 배포 중 색인이
돌고 있으면 같은 일이 난다는 뜻이라 백로그로 옮길 값이 있다.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
빈 청크 코퍼스 → 상한 2.0으로 색인 → 위치 P에서 E1 스윕 → 이긴 상한으로 완주 → 샘플러 정지 → 인스턴스 정지까지 한 번에 돈다. 사람이 자는 동안 돌 것이라, 손으로 칠 때는 사람이 하던 두 판단을 스스로 해야 한다. 이긴 상한 고르기. 원본당 ms가 가장 작은 조건을 고르되, 측정이 성립하지 않은 행 (note가 있거나 원본당 ms가 빈 행)은 뺀다. 값이 같으면 낮은 상한이 이긴다 -- 같은 처리량이면 블라스트 반경이 작은 쪽이 낫다는 6-11의 판단을 그대로 따른다. 못 고르면 2.0으로 완주한다. 밤을 멈추는 것보다 느리게라도 끝내는 편이 낫다. 색인 종료 판정. 앱 로그의 "인덱싱 종료"가 1차이고, 청크가 STALL_MIN분간 안 느는 것이 2차다. 로그만 보면 앱이 조용히 죽었을 때 영원히 기다리고, 정체만 보면 재시작 직후의 건너뛰기 구간을 종료로 오인한다. 둘 다 필요하다. P는 25,000청크로 잡았다. 6-11이 잰 자리와 같게 두면 스윕의 절대값을 그 기록과 나란히 놓을 수 있다. 출발선을 강제한다. 남아 있는 청크가 있으면 TRUNCATE한다. E8이 보려는 것이 빈 코퍼스 쪽 끝이라, 연습 주행이 남긴 3,273개를 그대로 두면 곡선의 왼쪽이 이미 없다. 성공이든 실패든 인스턴스를 정지한다. 정지는 EBS를 지우지 않으므로 코퍼스와 결과는 남는다. 실패했는데 켜둔 채 자면 요금만 는다 -- 계획 8-5가 적은 25만원짜리 실수가 바로 그것이다. 샘플러는 KILL이 아니라 TERM으로 끊는다. 그래야 구간별 표를 찍고 죽는다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
indexAll()은 예외를 잡아 커밋된 분량을 살린 뒤 "인덱싱 종료"를 찍는다. 완주한 실행과 죽은 실행이 같은 문자열로 끝난다는 뜻이다. 2026-08-13 밤 1 무인 측정이 여기서 무너졌다. 앱이 02:59:57에 뜨고 3초 뒤 03:00 정각 cron이 겹쳐 들어와 uk_chunk_source 유니크 충돌로 죽었는데, 그 실행이 찍은 "인덱싱 종료"를 스크립트가 완주로 읽고 60초 뒤 인스턴스를 껐다. 그때 다른 스레드는 아직 색인 중이었고, 코퍼스는 165,016원본 중 21,800원본(13%)에서 멈췄다. STATUS에는 OK가 남았다. 로그 문자열만 가르면 거짓말이 절반 남는다. int(신규 임베딩 수)를 돌려주는 한 호출자는 완주를 판정할 수 없다 -- 중단돼도 그때까지의 건수가 그대로 나오기 때문이다. 실제로 SeedIndexingRunner는 죽은 실행에도 "기동 직후 색인 완료"를 찍고 있었다. 그래서 반환형을 IndexRun(embedded, totalChunks, elapsedMs, abortReason)으로 바꿨다. 완주 여부는 completed()가 답하고, 로그도 "인덱싱 완주" / "인덱싱 중단 종료"로 갈린다. 사유는 무인 실행의 STATUS로 옮겨지므로, 메시지가 없는 예외는 예외 이름으로 대체한다. 동시 색인 자체 -- cron과 기동 러너가 같은 행을 집는 것 -- 는 아직 안 고쳤다. 이 커밋이 고치는 것은 "그것이 일어났을 때 알 수 있는가"다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
밤 1이 사람 대신 해야 했던 판단 둘이 모두 어긋났다. 완주 판정. "인덱싱 종료" 하나만 봤는데, 앱이 예외를 잡고도 같은 줄을 찍었다. 이제 "인덱싱 중단"이 보이면 그 자리에서 실패로 끝내고 사유를 STATUS에 옮긴다. 앱은 다음 주기에 이어서 진행할 수 있지만 측정은 그럴 수 없다 -- 코퍼스가 어디까지 찼는지 모르는 채로 이어 그린 E8 곡선은 읽을 수 없다. 끝난 이유(end_reason)도 함께 남긴다. 완주 로그와 정체 판정은 둘 다 rc=0으로 끝나지만 아침에 읽을 때 무게가 다르다. 승자 선택. 여섯 조건이 전부 75.0ms/원본으로 나왔고 sort -n | head -1이 첫 줄인 2.0을 집었다. "값이 같으면 낮은 상한이 이긴다"는 의도는 앞선 커밋에 이미 적어두었는데 구현이 그것을 하지 않았다 -- 조건 순서가 오름차순이라 우연히 낮은 값이 나왔을 뿐이고, 75.0 대 75.1처럼 노이즈만큼 갈리면 그대로 뒤집힌다. 그리고 무엇보다, 동률을 동률이라고 말하지 못하면 "상한은 이 워크로드의 제약이 아니다"라는 결론이 "2.0이 이겼다"로 보인다. 최고값의 +TIE_PCT%(기본 2) 안에 드는 조건을 동률로 보고 그중 가장 낮은 상한을 고른다. 동률 여부와 폭은 winner-why에 남는다. 곁가지로 awk의 n을 BEGIN에서 0으로 못박았다. 초기화 전에 첨자로 쓰면 첫 행이 ms[""]에 들어가고, 루프가 도는 ms[0]은 값이 없는 유령 행(0ms)이 된다. 그 유령은 항상 동률대에 들어가므로 동률 개수가 하나 부풀고 승자도 엉뚱한 줄에서 나온다. 실전 CSV에서 답이 맞았던 것은 우연이었다. 1구간 상한을 SEG1_CPUS로 뺐다. 승자가 그 값과 같으면 E8 두 구간의 조건이 바뀌지 않는다는 뜻이라, 곡선을 이어 그리기 전에 로그로 알린다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
결과를 쌓을 자리를 만들었다. 계획은 AWS-MEASUREMENT-PLAN.md에, 결과는 밤별로 AWS-MEASUREMENT-RESULTS.md에 남긴다. 로컬 실측(SERVICE-SCALE-ASSUMPTIONS 6절)은 호스트가 달라 절대값을 나란히 둘 수 없는 자리가 있으므로 대체하지 않는다. E1 -- 결론이 났다. 상한 2.0에서 6.0까지 여섯 조건이 전부 62.6ms/청크, 75.0ms/원본으로 같았다. 계획 E1의 성공 판정이 "상한을 올려도 원본당 ms가 안 변하면 병목이 TEI가 아니다"였고 그 가지로 떨어졌다. 근거 넷을 함께 적었다: 조건마다 cgroup cpu.max를 직접 읽었고, 통제군(마지막 2.0)이 첫 줄과 일치했고, 조건마다 같은 문서를 다시 색인했고, 앱 로그의 조건별 색인 지속 시간이 321.4~322.6초인데 넣은 양이 여섯 번 모두 4,907건이었다. 막힘의 배분이 결론을 한 겹 더 받친다. 2.0에서 TEI가 8.0% 막히는데도 처리량이 안 깎였고 app은 0.1%, db는 1.8%다. 어느 컨테이너도 CPU에 굶고 있지 않다. 병목은 CPU 포화가 아니라 PAGE_SIZE=100짜리 페이지 한 장의 약 7.5초 직렬 지연이다. 한계도 적었다. 조건이 2.0에서 위로만 갔다. 상한을 얼마나 줄 것인가가 원래 질문이라면 관심 구간은 아래쪽이고, 이 밤은 거기에 답하지 않는다. E8 -- 0에서 25,815청크까지 68표본이 75.0ms/원본으로 평평했다. HNSW 인덱스는 V1이 빈 테이블에 미리 만들므로 삽입마다 그래프를 갱신하는 비용이 이 값에 이미 들어 있다. 다만 26,285청크면 벡터와 그래프를 합쳐 50MB 안팎(추정)이라 shared_buffers 128MB 안에 전부 앉는다. 전량 색인이면 300MB를 넘겨 버퍼 밖으로 나가므로, E8이 보려던 열화는 멈춘 지점 이후에 시작된다. 6-11의 2.4배는 코퍼스 크기보다 호스트 경합 쪽이 유력해졌지만 확정은 완주 측정을 기다린다. 밤을 13%에서 끝낸 경위와 타임라인, 그리고 이 밤이 남긴 수정 넷(완료)과 하나(미완: 색인 상호배제)를 함께 적었다. 산출물은 measure-aws/에 그대로 넣었다 -- 표는 요약이고 원본은 CSV다. 스윕 구간의 "임베딩 서버 호출에 실패했습니다"가 실패가 아니라 스크립트가 앱을 정상 정지시킨 흔적이라는 것도 적어두었다. 다음 사람이 그 줄에서 멈추지 않도록. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
<hidden_range_assignment> 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 2📝 Generate docstrings 💡
⚔️ Resolve merge conflicts 💡
🧪 Generate unit tests (beta)
Comment |
compose.gc.yaml은 측정용 오버레이이고 이미 RAG_INDEXING_ON_STARTUP=true로 색인 트리거를 기동 러너에 맡기고 있다. 그러면 cron 배치(매일 03:00)는 필요가 없는데 켜져 있었다. 밤 1이 그 창에 걸렸다. 앱이 02:59:57에 뜨고 3초 뒤 정각 배치가 들어와 색인기 둘이 같은 행을 집었고 uk_chunk_source 유니크 충돌로 죽었다. 스윕은 조건마다 앱을 내렸다 올리므로 새벽 어느 시각에든 이 창이 열린다. 상호배제가 없는 것은 앱의 결함이고 따로 고쳐야 한다. 다만 그것을 고친 뒤에도 측정 중에 배치가 끼어드는 것 자체가 오염이라, 여기서 끄는 것은 우회가 아니라 조건이다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
완주 구간에는 중단 감지를 넣었는데 P까지 기다리는 1구간에는 없었다. 거기서 색인이 죽으면 청크가 안 늘 뿐이라 스크립트는 아무것도 모르고 MAX_INDEX_HOURS를 다 쓴다. 밤 하나가 그대로 날아가고, 아침에 남는 것은 "P에 도달하기 전에 8시간을 넘겼다"라는 사유뿐이라 무엇이 죽였는지도 모른다. 같은 신호를 같은 방식으로 본다. 앱 기동 직전 시각을 잡아두고 그 뒤의 로그만 훑는다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
인덱스의 모드가 100644이었다. 같은 디렉터리의 다른 셋(aws-bootstrap, e1 스윕, e8 샘플러)은 100755인데 이것만 빠져 있었다. 밤 1에서 안 드러난 이유는 그때의 전달 경로가 작업 트리를 그대로 복사해 로컬 권한을 들고 갔기 때문이다. git archive로 보내자 그 자리에서 죽었다. nohup: failed to run command './scripts/night1-run.sh': Permission denied b50665b가 gradlew에서 고친 것과 같은 종류다. 그때 "전달 경로가 tar나 zip이면 모드를 다시 잃을 수 있다"고 적어두고, 정작 그 뒤에 추가한 스크립트에 모드를 안 붙였다. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…lidation fix: 안전교육 이수 엔드포인트가 호출될 때마다 500이었다
There was a problem hiding this comment.
Actionable comments posted: 80
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/main/java/com/DOCKin/ai/service/FastApiService.java (1)
124-153: 🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy liftFastAPI 호출을 트랜잭션 밖으로 이동하십시오.
saveTranslateLog는 Line 124에서 트랜잭션을 시작합니다. Line 127의 조회 후 Line 141-153에서 외부 HTTP 응답을 동기 대기합니다. FastAPI가 느리거나 중단되면 DB 커넥션이 응답 대기 동안 점유됩니다. 커넥션 풀이 고갈되면 다른 요청도 DB 작업을 완료하지 못합니다.외부 번역을 먼저 완료하십시오. 번역 결과의 조회·갱신·저장만 별도
@Transactional메서드에서 처리하십시오.fastApiWebClient의 연결 및 응답 타임아웃도 명시적으로 제한하십시오.As per path instructions,
**/*.java: “트랜잭션 안에서 외부 HTTP를 호출하면 반드시 지적한다. 타임아웃이 없으면 커넥션 풀이 마르고 서비스 전체가 멎는다.”🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/main/java/com/DOCKin/ai/service/FastApiService.java` around lines 124 - 153, saveTranslateLog에서 WorkLog 조회와 FastAPI 호출이 시작되기 전에 트랜잭션을 열지 않도록 외부 번역 흐름을 분리하십시오. titleMono/contentMono의 번역 요청을 먼저 완료한 뒤, 번역 결과를 적용하고 조회·갱신·저장하는 부분만 별도 public `@Transactional` 메서드로 이동해 호출하십시오. fastApiWebClient 설정 또는 해당 요청 체인에 연결 및 응답 타임아웃을 명시적으로 추가하십시오.Source: Path instructions
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/ci.yml:
- Around line 47-86: Update the test job’s actions/checkout, actions/setup-java,
gradle/actions/setup-gradle, and actions/upload-artifact references to
40-character immutable commit SHAs, retaining version comments for readability.
Add test-job permissions with contents: read, and configure actions/checkout
with persist-credentials: false so the token is not stored locally.
In `@docs/adr/0001-attendance-clockin-concurrency.md`:
- Line 96: Revise the lock guidance in the attendance clock-in concurrency ADR
so the lock TTL exceeds the maximum protected doClockIn and database commit
duration, or explicitly require a watchdog to renew the lease while work is
active. Replace the incorrect 3-second example and state whether any proposed
duration is based on measurements or an assumption.
In `@docs/adr/0002-performance-improvement-backlog.md`:
- Around line 41-43: Update the “JVM GC/힙 튜닝” section to identify the active
heap limit as -Xmx512M, reflecting the Docker ENTRYPOINT behavior that does not
expand JAVA_OPTS. Revise the supporting explanation of container memory headroom
and GC-tuning rationale to match the effective configuration, and mark any
unverified configuration claims with “[미검증]”.
In `@docs/adr/0003-search-domain-expansion.md`:
- Line 10: Update the assumptions in the ADR, including the traffic and
infrastructure statements near “DOCKin” and the 0.02 TPS and 100,000-row chunk
settings near the relevant roadmap entry, to clearly distinguish measured
results from planning assumptions. Add supporting evidence where available;
otherwise mark unverified values with “[미검증]” and assumptions requiring
validation with “[실측 필요]”.
In `@docs/adr/0005-attendance-portfolio-gap-analysis.md`:
- Around line 91-95: Update the CI/CD assessment in the “반대로 뒤집힌 것” section to
reflect the current .github/workflows/ci.yml workflow. Verify its actual
execution status, distinguish between a missing workflow and an existing
workflow that fails, and revise the supporting evidence accordingly. Also flag
any conflicting test-count figures elsewhere in the documentation.
In `@docs/adr/0006-rag-vector-search.md`:
- Around line 3-30: 문서 머리말과 2절 결정 요약을 현재 구현 상태에 맞게 갱신하십시오. 8-2와 8-3에 기록된
PostgreSQL·pgvector·HNSW·SQL-side vector search 완료 상태를 현재 상태로 먼저 요약하고, MySQL
브루트포스 설계와 ANN 도입 판단은 Phase 1의 과거 결정으로 별도 구분하십시오. 다른 문서와 수치가 다르면 해당 불일치를 식별해
정정하거나 출처를 명시하십시오.
In `@docs/AWS-MEASUREMENT-PLAN.md`:
- Around line 419-421: 문서의 해당 비용 비교에서 최악 $23와 4 vCPU 최선 $13의 차액을 1,400원/USD 환율
기준인 약 1.4만원으로 수정하고, 이어지는 E1 정산 및 m7i.xlarge 전환 판단의 근거도 이 차액과 일치하도록 갱신하십시오.
In `@docs/AWS-MEASUREMENT-RESULTS.md`:
- Around line 60-61: docs/AWS-MEASUREMENT-RESULTS.md의 E1 처리량 수치를 하나의 권위 있는 출처와
값으로 통일하십시오. 표의 4,790 청크와 본문의 4,907건 중 실제 기준을 선택하고, 다른 값이 별도 지표라면 단위와 정의를 명확히
구분하십시오. 처리량 및 “5%면 200건” 계산이 동일한 기준으로 재현되도록 관련 설명과 산출값을 함께 수정하십시오.
In `@docs/db/rebuild-hnsw-index.sql`:
- Around line 52-53: Update the rebuild procedure around
idx_chunk_embedding_hnsw so rerunning it actually recreates the HNSW index:
explicitly drop the existing index before creating it, or remove IF NOT EXISTS
so an existing index causes an immediate failure. Ensure the chosen behavior
makes regeneration failures visible to the operator.
- Around line 63-64: Update the Korean comments near the HNSW index definition
to reflect that RetrievalService and DocumentChunkRepository already delegate
similarity ordering through ORDER BY embedding <=>, rather than claiming the
index is unused. Describe the current query and execution-plan verification
procedure, and mark any unverified configuration with the “[미검증]” label.
In `@docs/PORTFOLIO-ROADMAP.md`:
- Around line 134-137: Update the performance multiplier in the “꼭지 1” roadmap
entry to match the stated 5,165ms-to-57ms comparison: use approximately 91x, or
explicitly identify a different baseline if retaining 70x.
In `@docs/PROJECT-SCOPE.md`:
- Around line 45-46: Update the Flyway migration description in
docs/PROJECT-SCOPE.md to identify V2 as the existing-table baseline and include
V3__work_log_child_fk_indexes.sql, V4__users_child_fk_indexes.sql, and
V5__work_logs_user_id_not_null.sql as subsequent migrations. Ensure the text no
longer implies that the migration chain ends at V2.
In `@docs/SERVICE-SCALE-ASSUMPTIONS.md`:
- Around line 114-125: 폐기된 Phase 1 가정 수치가 현재 계획의 근거로 보이지 않도록 출처와 상태를 정리하십시오.
docs/SERVICE-SCALE-ASSUMPTIONS.md 114-125에서는 12 ms/건 인덱싱 표와 “병목이 아니다” 결론을 과거
추정으로 표시하고 6-8의 실제 파이프라인 측정값을 연결하며, 170-171에서는 Xmx400M을 제거하거나 당시 미적용 가정으로 명시하십시오.
docs/adr/0006-rag-vector-search.md 114-120에서는 12 ms/건 표를 현재 트리거 근거에서 제거하고 후속 실측
또는 역사적 기준선으로 표시하며, 491-498에서는 원본당 2~3청크를 1.22 실측값과 한계로 갱신하십시오.
docs/adr/0007-corpus-retention.md 14에서는 1.22 청크/원본 실측에 따라 200~300만 청크 용량 계획을
재계산하고 외삽값에는 [실측 필요]를 유지하십시오.
In `@measure-aws/night1-20260812T164952/STATUS`:
- Around line 1-2: Update the STATUS artifact to match the current
scripts/night1-run.sh success format by including OK, chunks=26285, and an
end_reason=... line. If this reflects a pre-change run, explicitly mark that
fact; otherwise verify the stop reason and regenerate the artifact, preserving
25,815 as the measurement point and 26,285 as the stopping chunk count.
In `@nginx/conf.d/default.conf`:
- Around line 66-68: Align the proxy buffer configuration around
proxy_buffer_size, proxy_buffers, and proxy_busy_buffers_size with the
dockin-nginx 100M memory limit: either increase the container limit to support
the current buffers or reduce the buffer count/size to retain 100M. Document the
rationale beside the chosen values, using measured evidence when available or
“[실측 필요]” when based on an assumption.
- Around line 94-98: Update the dockin-app service configuration in compose.yaml
to remove the host-published 8080:8080 port and retain only expose: "8080",
ensuring clients access the app through nginx’s /ws proxy path.
In `@scripts/aws-bootstrap.sh`:
- Around line 120-124: Update the variable-extraction regex in the loop over
compose.yaml and compose.gc.yaml so it matches only references ending
immediately after the variable name, excluding forms with default values such as
${VAR:-default}; preserve the existing MISSING collection and die behavior for
required variables.
- Around line 44-47: Update the BOOTSTRAP_REEXEC assignment immediately before
exec sg docker in the bootstrap re-execution flow to preserve the one-time guard
by setting it to 1 instead of 0; leave the existing re-execution command
unchanged.
- Around line 57-60: Update the Docker Compose download logic in the bootstrap
script to derive the release architecture from uname -m instead of hardcoding
x86_64. Map the host architecture to the archive naming expected by the Compose
release URL, including arm64 instances, and use that value when constructing the
curl URL while preserving the existing installation path and permissions.
In `@scripts/e1-tei-cpu-sweep.sh`:
- Around line 289-302: Update the embedding wait logic around the `waited` loop
to use an explicit flag for the “embedding never started” timeout outcome, set
it only when the timeout branch records the CSV row, and use that flag for
cleanup and `continue` after the loop. Do not infer failure from `waited >=
EMBED_WAIT_MAX`, so a successful start at the boundary still proceeds to open
the window while every timeout retains its CSV record.
- Around line 222-228: Update the container CPU aggregation around OTHERS to
detect any service whose docker inspect NanoCpus value is 0, treating it as an
unlimited CPU container rather than adding 0. If any unlimited container is
found, skip the numeric saturation-boundary and OVER calculations, record the
boundary/oversubscribed results as n/a (or stop the sweep explicitly), and
ensure the CSV cannot report a misleading “no”.
- Around line 359-360: Update the corpus-position validation around AFTER and
CORPUS_AT_START to fail the measurement sweep with die, matching the cgroup
validation behavior near the existing check, instead of only emitting say and
continuing. Preserve an environment-variable override for flow-checking runs so
those runs may continue intentionally; when override behavior permits
continuation, record the rollback mismatch in the CSV note field.
In `@scripts/e8-index-sampler.sh`:
- Around line 112-118: Update the ROW assignment and error handling in the
sampling loop around psql_q so a database query failure is captured without
triggering set -e termination, allowing the existing db-unreachable CSV
recording branch to execute. Preserve the current successful-query processing
and failure-row fields.
In `@src/main/java/com/DOCKin/absence/service/AbsenceRequestService.java`:
- Around line 47-53: Update requirePendingRequest() to load the AbsenceRequest
with pessimistic write locking (or apply equivalent `@Version` conflict handling)
so concurrent approveRequest() calls cannot process the same PENDING request
twice; move S3 putObject out of createRequest()’s database transaction and
configure finite connection and socket timeouts on the S3 client.
Apply the same fix in
`@src/main/java/com/DOCKin/absence/controller/AbsenceAdminController.java` around
lines 41 - 58.
- Around line 56-68: Move the S3 upload out of the `@Transactional` createRequest
method while preserving the existing validation and persistence flow. Add
compensating cleanup to delete the uploaded object when the subsequent database
save fails, rethrowing the original failure and retaining the current null
behavior when no document is provided.
Apply the same fix in
`@src/main/java/com/DOCKin/absence/controller/AbsenceAdminController.java` around
lines 34 - 40: 같은 요청 생성 경로의 S3 업로드 호출을 확인하는 위치입니다.
In `@src/main/java/com/DOCKin/ai/controller/AiController.java`:
- Around line 55-72: Update FastApiService.chatBotFromSpringToFastApi so
connection and response timeout exceptions from the WebClient call are converted
to CHATBOT_NOT_WORK, alongside existing HTTP status error handling. Ensure these
transport exceptions do not reach the global catch-all as INTERNAL_SERVER_ERROR,
while preserving other exception behavior.
In `@src/main/java/com/DOCKin/ai/model/TranslateLog.java`:
- Around line 29-45: 기존 translate_logs 테이블과 데이터를 새 work_log_translations 구조로
이전하는 별도 Flyway 마이그레이션을 추가하거나, 이전이 완료되지 않은 경우를 명시적으로 차단하는 사전 조건 검사를 추가하십시오.
src/main/java/com/DOCKin/ai/model/TranslateLog.java:29-45,
src/main/java/com/DOCKin/ai/model/ChatLog.java:40-45,
src/main/java/com/DOCKin/attendance/model/Attendance.java:11-29의 신규 매핑과 기존 데이터
호환성을 확인하고, 세 위치에는 직접 변경이 필요 없으며 마이그레이션 또는 검사로 해결하십시오.
In `@src/main/java/com/DOCKin/ai/service/FastApiService.java`:
- Around line 171-186: Make the find/update-or-save flow in FastApiService
atomic for concurrent retranslations. Use PostgreSQL INSERT ... ON CONFLICT
(log_id, language_code) DO UPDATE, or serialize the lookup and mutation by
locking the related WorkLog row within one transaction; preserve updating the
existing translation and inserting a new one with the current request values
without allowing duplicate-key failures.
In `@src/main/java/com/DOCKin/attendance/service/AttendanceService.java`:
- Around line 56-61: Update the lock acquisition flow around RLock.tryLock in
AttendanceService so InterruptedException is caught separately from Redis
failures, the thread interrupt status is restored, and the request returns an
application error without calling doClockIn. Keep the existing DB-constraint
fallback only for non-interruption lock failures.
In `@src/main/java/com/DOCKin/checklist/dto/ChecklistItemUpdateRequestDto.java`:
- Around line 17-22: Update ChecklistItemUpdateRequestDto at
src/main/java/com/DOCKin/checklist/dto/ChecklistItemUpdateRequestDto.java:17-22
by adding `@Size`(max = 255) to content, and update ChecklistUpdateRequestDto at
src/main/java/com/DOCKin/checklist/dto/ChecklistUpdateRequestDto.java:17-21 by
adding `@Size`(max = 100) to title. Preserve nullable optional-field behavior by
using `@Size` rather than a non-null constraint.
In `@src/main/java/com/DOCKin/checklist/service/ChecklistService.java`:
- Around line 65-67: Update createChecklist and updateChecklist to explicitly
flush after persisting changes, catch DataIntegrityViolationException caused by
the uk_checklist_equipment_phase constraint, and convert it to
BusinessException(ErrorCode.CHECKLIST_ALREADY_EXISTS). Preserve the existing
duplicate pre-check while ensuring concurrent requests receive the same business
error.
In `@src/main/java/com/DOCKin/global/config/WebClientConfig.java`:
- Around line 39-40: WebClientConfig 주석의 임베딩 성능 수치를 application.properties의 정정된
실측값과 일치시키십시오. “배치 32건, 건당 약 12ms, 정상 응답 1초 이내”라는 설명을 청크 길이에 따른 건당 46~153ms 및 배치
한 번이 20초를 초과할 수 있다는 내용으로 수정하고, 짧은 타임아웃을 유도하는 표현을 제거하십시오.
- Line 4: Update the HttpHeaders import used by WebClientConfig to Spring’s
HttpHeaders class instead of org.apache.http.HttpHeaders, leaving the existing
header configuration behavior unchanged.
In `@src/main/java/com/DOCKin/global/logging/TraceIdFilter.java`:
- Around line 50-52: TraceIdFilter에서 컨트롤러 실행 후 최종 MDC에 저장된 TraceId를 기준으로 응답
X-Trace-Id 헤더를 갱신하십시오. TraceId.override()가 호출된 AI 경로에서는 기존 traceId 대신 최종 MDC 값을
사용하고, 본문 traceId를 우선하는 기존 동작과 일치시키십시오.
In `@src/main/java/com/DOCKin/global/security/config/SecurityPathConfig.java`:
- Around line 15-19: Initialize the whitelist field in SecurityPathConfig with
an empty list so it is never null when getWhiteListArray() is called. Preserve
the existing array conversion behavior and use the suggested ArrayList import or
an equivalent empty-list initialization.
In `@src/main/java/com/DOCKin/member/model/Member.java`:
- Around line 66-71: Update Member.useLeaveDays to reject non-positive days
before applying the balance deduction, while preserving the existing
insufficient-balance check for valid positive values. Ensure invalid input
raises the appropriate business validation exception and cannot increase
remainingLeaveDays.
- Around line 45-52: 기존 users 테이블을 보완하는 별도 Flyway 증분 마이그레이션을 추가하십시오. workShift와
remainingLeaveDays 컬럼이 없을 때 추가하고, 기존 remainingLeaveDays 행을 15로 채운 후 해당 컬럼에만 NOT
NULL 제약을 적용하십시오. workShift에는 NOT NULL 제약을 추가하지 말고,
V2__baseline_existing_tables.sql의 스키마와 엔티티 매핑에 맞는 컬럼명을 사용하십시오.
In `@src/main/java/com/DOCKin/member/repository/MemberRepository.java`:
- Around line 36-49: Update the Javadoc in MemberRepository to document the
current PostgreSQL lock-timeout behavior: reference the configured
lock_timeout=5s and explain that SET LOCAL lock_timeout applies within the
transaction. Mark the existing MySQL measurement as historical, and remove or
replace references to innodb_lock_wait_timeout, --innodb-lock-wait-timeout=5,
and the claim that no application path exists.
In `@src/main/java/com/DOCKin/member/service/MemberService.java`:
- Line 83: Remove the silent MORNING fallback from the MemberService member
construction flow. Require MemberRequestDto.workShift via `@NotNull` so clients
must provide it, or centralize the default on the Member entity field and pass
dto.getWorkShift() directly without service branching; choose one approach and
ensure missing values cannot be misclassified as MORNING.
In `@src/main/java/com/DOCKin/rag/service/ChunkIndexWriter.java`:
- Around line 63-71: Update writePage and collectPending to batch existing-chunk
retrieval: add the repository method
findBySourceTypeAndSourceIdInAndEmbeddingModel, query once using the page’s
source IDs, group results by sourceId, and pass each target’s existing chunks
into collectPending instead of querying per target.
- Around line 50-57: writePage currently holds its transaction while flush
performs multiple external embedding calls, allowing DB connection usage to grow
with page size. Separate preparation/embedding from transactional persistence,
or at minimum enforce a page-level total embedding time limit across all calls
in flush so the transaction cannot exceed a finite bound; preserve the existing
batch processing and storage behavior.
In `@src/main/java/com/DOCKin/rag/service/EmbeddingClient.java`:
- Around line 47-48: Validate the injected batchSize configuration during
application startup, rejecting any value less than or equal to zero before
embedPassages can run. Update the EmbeddingClient initialization/configuration
path around the batchSize field, while preserving positive batch sizes for the
existing embedPassages batching loop.
- Around line 110-120: Update the catch block in the embedding client method to
pass the caught exception e as the final argument to log.error, while preserving
the existing exception class and message fields. Keep the BusinessException
behavior unchanged unless an existing cause-aware constructor is available,
ensuring the original stack trace and nested cause details are recorded.
In `@src/main/java/com/DOCKin/rag/service/IndexingService.java`:
- Around line 144-146: Move the countByEmbeddingModel call into the existing
protected error-handling flow in the indexing method so a count failure cannot
prevent IndexRun creation. Preserve progress.embedded and abortReason, and
provide a safe fallback for total when the count lookup fails so the method
still returns its completion result.
- Around line 90-97: scheduledIndexing에서 indexAll()의 반환값을 버리지 말고 완료 여부를 지역 변수로
받아 명시적으로 다루세요. 인덱싱 비활성화 시 기존 조기 반환은 유지하고, 반환값을 로그 문자열에 의존하지 않는 후속 메트릭·알림 처리 지점으로
사용할 수 있게 하세요.
- Around line 225-238: Update toTarget(TranslateLog) so translations with a null
work-log source are excluded from indexing and the orphan condition is recorded
through the existing logging mechanism. Ensure no IndexTarget is created for
that case, then simplify ownerUserId creation to use the non-null source
directly while preserving Visibility.OWNER for valid translations.
In `@src/main/java/com/DOCKin/rag/service/RagChatService.java`:
- Around line 58-62: Update the FastAPI chat response flow used by
fastApiService.chatBotFromSpringToFastApi so an empty response body is converted
to the established domain exception via switchIfEmpty(Mono.error(...)) before
block() returns. Preserve the existing blocking behavior and ensure saveChatLog
receives only a non-null response.
In `@src/main/java/com/DOCKin/rag/service/RetrievalService.java`:
- Around line 134-149: Update RetrievalService.search so
embeddingClient.embedQuery is executed before entering any transactional
database work, while preserving the existing validation and dimension-mismatch
handling. Move the vector-search/database portion into a separate transactional
bean or method invoked through a transaction proxy; do not rely on same-bean
self-invocation, and ensure only the DB lookup holds the read-only transaction.
Apply the same fix in
`@src/main/java/com/DOCKin/rag/service/RetrievalService.java` around lines 158 -
162: 동일한 외부 임베딩 호출의 트랜잭션 내 실행과 커넥션 점유 위험을 다룹니다.
In `@src/main/java/com/DOCKin/rag/service/SeedIndexingRunner.java`:
- Around line 29-33: Update the Javadoc condition above the SeedIndexingRunner
declaration to document the actual `@Profile`("seed") requirement, replacing the
incorrect local profile reference while preserving the startup indexing property
condition.
In `@src/main/java/com/DOCKin/worklog/controller/WorkLogsController.java`:
- Line 86: 목록 API의 createdAt 및 logId 내림차순 정렬을 지원하도록 work_logs에 두 컬럼의 복합 인덱스를
추가하십시오. 기존 마이그레이션을 수정하지 말고 새 마이그레이션 버전 파일을 생성하며, WorkLogsController의 pageable 정렬
순서와 인덱스 컬럼 순서를 일치시키십시오.
In `@src/main/java/com/DOCKin/worklog/service/WorkLogsService.java`:
- Around line 40-41: WorkLogsService.createWorklog의 `@Transactional` 범위에서 S3 업로드와
sttService.processStt(...).block()을 분리하고, 외부 작업을 트랜잭션 없이 먼저 수행한 뒤 DB 저장 부분만 별도의
짧은 트랜잭션으로 실행하도록 변경하십시오. 동일하게 적용되는 73-75행의 트랜잭션도 조정하고, 업로드 실패 시 기존 정책에 맞는 정리 또는
재시도 처리를 추가하십시오.
In `@src/main/resources/application.properties`:
- Around line 166-180: Update EmbeddingClient to use the configured
external-api.embedding.response-timeout-ms value for its blocking timeout
instead of hardcoding Duration.ofSeconds(60). Reuse the existing WebClientConfig
timeout property or remove the duplicate blocking timeout while preserving the
configured response-timeout behavior.
In `@src/main/resources/db/migration/V4__users_child_fk_indexes.sql`:
- Around line 25-26: V4의 work_logs 인덱스 생성이 운영 중 쓰기를 차단하지 않도록 처리하십시오. V4가 아직 운영
DB에 적용되지 않은 경우 Flyway 비트랜잭션 실행과 CREATE INDEX CONCURRENTLY를 사용하도록 해당 마이그레이션을
변경하고, 이미 적용된 환경에서는 V4를 수정하지 말고 후속 마이그레이션 또는 유지보수 작업으로 분리하십시오.
In `@src/main/resources/templates/chat_test.html`:
- Line 207: Update the input element identified by inputUserId to provide an
accessible name by associating it with a descriptive label or adding an
aria-label that identifies the field as an employee ID.
- Around line 47-49: Update the URL constants near API_BASE_URL, SERVER_URL, and
SOCKET_URL to use secure HTTPS/WSS connections through an authenticated domain,
preferably deriving the HTTP base from window.location.origin and constructing
the WebSocket URL with the matching secure protocol. Ensure the JWT and chat
traffic are never sent over plaintext HTTP or WS.
In `@src/test/java/com/DOCKin/absence/service/AbsenceRequestServiceTest.java`:
- Around line 189-205: Update approveRequest_sufficientLeaveDays_deductsBalance
to capture and verify the AbsenceApprovedEvent published through eventPublisher
when absenceRequestService.approveRequest succeeds. Assert that one event is
published and that its request or identifying data matches the approved request,
using the actual event accessors; retain the existing balance and response
assertions.
In `@src/test/java/com/DOCKin/absence/service/LeaveBalanceConcurrencyTest.java`:
- Around line 75-77: Remove the Assumptions.abort skip path from both test
methods, including the SQLException handling around the affected code, so
container connection failures fail the tests instead of being reported as
skipped; also remove the now-unused Assumptions import and update the Javadoc
that incorrectly claims there is no conditional skip path.
- Around line 160-161: Update the worker exception handling in
LeaveBalanceConcurrencyTest to collect exceptions in a thread-safe workerErrors
list instead of ignoring them. After the interleaveFailed validation, report the
collected worker errors so lock timeouts remain distinguishable from unexpected
database or connection failures.
In `@src/test/java/com/DOCKin/ai/service/TranslateParallelBenchmarkTest.java`:
- Around line 53-56: Update TranslateParallelBenchmarkTest by storing the
ExecutorService created for server.setExecutor(Executors.newFixedThreadPool(4))
in a field, then call shutdownNow() on that field from tearDown() alongside
server.stop(0).
- Around line 81-84: Update the timing assertion in
TranslateParallelBenchmarkTest so it uses a stable absolute performance
threshold based on the fixed 300ms stub delay, rather than requiring parallelMs
to be strictly less than sequentialMs. Preserve the benchmark output and ensure
the assertion allows normal CI scheduling variance while still detecting an
unexpectedly slow Mono.zip execution.
In `@src/test/java/com/DOCKin/attendance/service/AttendanceBatchServiceTest.java`:
- Around line 138-149: The tests currently cover only direct
markAbsentFor(LocalDate) calls, leaving markYesterdayAbsentees() and its
fixed-clock date calculation unverified. Add a test in the
AttendanceBatchServiceTest flow that enables batchEnabled, calls
markYesterdayAbsentees() on a service created by service(LocalDate), and
verifies markAbsentFor receives the day before the fixed date.
In `@src/test/java/com/DOCKin/global/config/ActuatorEndpointTest.java`:
- Around line 93-98: Update the Gradle test-task configuration to depend on
bootBuildInfo, ensuring build-info.properties is generated before tests run.
Keep the existing infoExposesBuildInformation test unchanged so BuildProperties
and the /actuator/info build section are available during execution.
In `@src/test/java/com/DOCKin/global/config/LocalMigrationDriftTest.java`:
- Around line 199-203: Update the .env parsing logic around the values map to
remove an optional export prefix from each entry and strip matching surrounding
single or double quotes from the parsed value before storing it. Preserve
unquoted values and internal quote characters, and ensure quoted DB_PASSWORD
values are passed to the database connection without the quote characters.
In `@src/test/java/com/DOCKin/global/config/SchemaValidationTest.java`:
- Around line 20-23: Update the class-level documentation in
SchemaValidationTest so it no longer claims that only V1/V2 create the schema;
describe the full migration set generically or state the current V1–V5 range
without implying a fixed future limit.
In `@src/test/java/com/DOCKin/global/web/PageableSortDefaultTest.java`:
- Around line 112-119: Update hasDefaultSort in PageableSortDefaultTest so it
does not treat merely having a sort attribute as guaranteeing a total order;
validate that the endpoint’s final sort key is unique, or explicitly narrow the
test/documentation contract to checking only sort presence. Also verify that the
declared PageableDefault or SortDefault annotation is actually applied by the
endpoint configuration.
In `@src/test/java/com/DOCKin/member/repository/LockTimeoutVerificationTest.java`:
- Around line 88-120: Update LockTimeoutVerificationTest to validate
PostgreSQL’s configured lock_timeout before attempting lock contention: use a
separate connection to execute SHOW lock_timeout, parse the result, and assert
it equals SERVER_TIMEOUT_MS, treating 0 or missing configuration as an immediate
failure. Strengthen the serverMs assertions to require the measured wait to be
approximately SERVER_TIMEOUT_MS within an appropriate tolerance, preventing long
or indefinite waits while preserving the existing local-timeout comparison.
In `@src/test/java/com/DOCKin/rag/repository/FlywayMigrationTest.java`:
- Around line 150-164: Update the test method 멱등성 to load the actual V1
migration file and execute its contents against the scratch database instead of
duplicating SQL statements in the test. Preserve the existing assertDoesNotThrow
behavior and resolve the migration path using the repository’s actual file
location, so changes to V1 are directly covered.
In
`@src/test/java/com/DOCKin/rag/repository/HibernateBatchInsertVerificationTest.java`:
- Around line 124-150: Update HibernateBatchInsertVerificationTest so
resetStatements(Connection) propagates pg_stat_statements reset failures instead
of ignoring SQLException, causing the test to fail when observation is
unavailable. In the verification assertions, require sequenceCalls > 0 and
enforce the same expectedCalls * 2 upper bound used by the diagnostic output,
replacing the permissive TOTAL / 2 threshold.
In `@src/test/java/com/DOCKin/rag/repository/VectorTypeMappingTest.java`:
- Around line 87-96: Update the 차원_불일치_거부 test to assert the specific Spring
DataAccessException expected from the database rejection instead of generic
Exception, capture it with assertThrows, and verify rootMessage contains
“vector” or dimension-related text. Add the required JUnit and Spring DAO static
imports and the root-cause message helper.
In `@src/test/java/com/DOCKin/rag/service/AnnRecallMeasurementTest.java`:
- Around line 141-156: Unify recall calculation in the query loop with the
existing avgRecall definition by using the exact nearest-neighbor result count
as the denominator instead of TOP_K. Update the corresponding “recall@5”
output/header labeling if needed so it clearly reflects this recall definition,
while preserving the existing intersection count and reporting flow.
- Around line 119-168: Restrict the try/catch in AnnRecallMeasurementTest.java
(119-168), including the 선택도_스윕 (191-228) and efSearch_스윕 (242-293) flows, to
DriverManager.getConnection so count, search, explainMs, and measurement
failures propagate; only connection failures should call Assumptions.abort.
Apply the same boundary in BruteForceSearchBenchmarkTest.java (54-81), allowing
prepareTable and measure failures to fail the test, and in
CorpusSeedGeneratorTest.java (130-157), allowing DELETE/INSERT failures to fail
rather than skip while retaining skip handling only for connection failures.
In `@src/test/java/com/DOCKin/rag/service/BruteForceSearchBenchmarkTest.java`:
- Around line 84-98: Update the comment above prepareTable to remove the
outdated VARBINARY and “identical to the real schema” claims, and state that
this benchmark reproduces only the Phase 1 baseline BYTEA path. Leave the CREATE
TABLE definition unchanged.
- Around line 61-81: Move the bench_document_chunks cleanup from the normal
execution path into a finally block surrounding the benchmark work in the test
method. Ensure dropTable(conn) is attempted when prepareTable or measure throws,
while preserving the existing SQLException handling and benchmark output
behavior.
In `@src/test/java/com/DOCKin/rag/service/CorpusSeedGeneratorTest.java`:
- Around line 793-801: Update the longForm Vietnamese text construction around
the affected sb.append calls so each complete format string, including its
concatenated literal, is formatted with months, similarCases, or leadWeeks as
appropriate; ensure no %d placeholders remain in generated text. Add the
requested assertion to the slot-correspondence test for both bodyVi() and
bodyKo().
- Around line 242-267: Update removePreviouslyGenerated to execute all three
DELETE statements within a single transaction: disable auto-commit, commit only
after every deletion succeeds, and roll back on any SQLException. Restore the
connection’s original auto-commit state before returning or propagating the
error, while preserving the existing Removed counts.
- Around line 469-477: Update the generated-key handling around the ResultSet
keys loop to consume all returned rows and verify the count exactly matches
compositions.size(). Do not use the positional index to associate keys with
compositions; map each generated key to its corresponding batch row using the
stable row identifier available in the generated-key result and retain the
correct owners, compositions, and dates association when creating
PendingTranslation.
In
`@src/test/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusValidationTest.java`:
- Around line 41-55: Update SafetyWatchStatusRequestDto.courseId to use `@NotNull`
instead of `@NotBlank`, matching its Integer type and preventing validation type
errors. In SafetyWatchStatusValidationTest, replace print-only handling with
assertions that null courseId produces a violation while a valid integer
produces none; keep the test failing on unexpected RuntimeException.
In `@src/test/java/com/DOCKin/worklog/service/WorkLogListBenchmarkTest.java`:
- Around line 568-577: Update the user deletion statement in cleanUp(Connection)
to match exactly the benchmark IDs generated by the test: the bch prefix
followed by four digits, rather than any user_id beginning with bch. Keep the
existing cleanup of benchmark-created users while excluding unrelated user
records.
- Around line 186-192: 분석벤치마크 테스트의 전체 실행을 감싸는 SQLException catch를 제거하고,
getConnection 단계에서 발생한 접속 실패만 Assumptions.abort로 처리하도록 연결 획득과 측정·정리 단계(load,
measureAll, createBenchIndexes, cleanUp)를 분리하십시오. 측정 또는 cleanUp에서 발생한
SQLException은 건너뛰지 말고 테스트 실패로 전파되게 하며, 기존 정리 동작은 유지하십시오.
In `@src/test/java/com/DOCKin/worklog/service/WorkLogListQueryCountTest.java`:
- Around line 159-177: Update the paginated measurement calls in the test,
including the all10, all20, other, and search cases, to use PageRequest.of with
Sort.by("logId"). Keep the measurement assertions unchanged and align these
requests with the ORDER BY log_id assumptions used by the later precondition
checks.
---
Outside diff comments:
In `@src/main/java/com/DOCKin/ai/service/FastApiService.java`:
- Around line 124-153: saveTranslateLog에서 WorkLog 조회와 FastAPI 호출이 시작되기 전에 트랜잭션을
열지 않도록 외부 번역 흐름을 분리하십시오. titleMono/contentMono의 번역 요청을 먼저 완료한 뒤, 번역 결과를 적용하고
조회·갱신·저장하는 부분만 별도 public `@Transactional` 메서드로 이동해 호출하십시오. fastApiWebClient 설정 또는
해당 요청 체인에 연결 및 응답 타임아웃을 명시적으로 추가하십시오.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 7bbf56ce-73db-4539-ab99-b39cf5ff4d07
⛔ Files ignored due to path filters (11)
app.jaris excluded by!**/*.jar,!**/*.jarmeasure-aws/app-night1.logis excluded by!**/*.logmeasure-aws/e1-20260812T164256/e1-sweep.csvis excluded by!**/*.csvmeasure-aws/e1-20260812T164256/run.logis excluded by!**/*.logmeasure-aws/e1-20260812T171704/e1-sweep.csvis excluded by!**/*.csvmeasure-aws/e1-20260812T171704/run.logis excluded by!**/*.logmeasure-aws/e8-20260812T164954/e8-samples.csvis excluded by!**/*.csvmeasure-aws/night1-20260812T164952/e1-sweep.csvis excluded by!**/*.csvmeasure-aws/night1-20260812T164952/sampler.logis excluded by!**/*.logmeasure-aws/night1-20260812T164952/sweep.logis excluded by!**/*.logmeasure-aws/night1.logis excluded by!**/*.log
📒 Files selected for processing (223)
.coderabbit.yaml.dockerignore.env.example.gitattributes.github/workflows/ci.yml.github/workflows/workflow.yml.gitignoreDockerfileREADME.mdbuild.gradlecompose.gc.yamlcompose.nocpu.yamlcompose.yamldocs/2026-07-04-checklist-domain-work-summary.mddocs/AWS-MEASUREMENT-PLAN.mddocs/AWS-MEASUREMENT-RESULTS.mddocs/JVM-GC-COLLECTORS.mddocs/PORTFOLIO-ROADMAP.mddocs/PROJECT-SCOPE.mddocs/SERVICE-SCALE-ASSUMPTIONS.mddocs/WORK-BACKLOG.mddocs/adr/0001-attendance-clockin-concurrency.mddocs/adr/0002-performance-improvement-backlog.mddocs/adr/0003-search-domain-expansion.mddocs/adr/0004-high-traffic-scaling-roadmap.mddocs/adr/0005-attendance-portfolio-gap-analysis.mddocs/adr/0006-rag-vector-search.mddocs/adr/0007-corpus-retention.mddocs/db/postgresql-schema.sqldocs/db/rebuild-hnsw-index.sqlgradlewinit.sqlmeasure-aws/e1-20260812T164256/env.mdmeasure-aws/e1-20260812T164256/raw/seq1-DOCKin-DB.txtmeasure-aws/e1-20260812T164256/raw/seq1-dockin-app.txtmeasure-aws/e1-20260812T164256/raw/seq1-dockin-embedding.txtmeasure-aws/e1-20260812T164256/raw/seq1-dockin-nginx.txtmeasure-aws/e1-20260812T164256/raw/seq1-dockin-redis.txtmeasure-aws/e1-20260812T164256/raw/seq2-DOCKin-DB.txtmeasure-aws/e1-20260812T164256/raw/seq2-dockin-app.txtmeasure-aws/e1-20260812T164256/raw/seq2-dockin-embedding.txtmeasure-aws/e1-20260812T164256/raw/seq2-dockin-nginx.txtmeasure-aws/e1-20260812T164256/raw/seq2-dockin-redis.txtmeasure-aws/e1-20260812T171704/env.mdmeasure-aws/e1-20260812T171704/raw/seq1-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq1-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq1-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq1-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq1-dockin-redis.txtmeasure-aws/e1-20260812T171704/raw/seq2-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq2-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq2-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq2-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq2-dockin-redis.txtmeasure-aws/e1-20260812T171704/raw/seq3-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq3-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq3-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq3-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq3-dockin-redis.txtmeasure-aws/e1-20260812T171704/raw/seq4-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq4-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq4-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq4-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq4-dockin-redis.txtmeasure-aws/e1-20260812T171704/raw/seq5-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq5-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq5-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq5-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq5-dockin-redis.txtmeasure-aws/e1-20260812T171704/raw/seq6-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq6-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq6-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq6-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq6-dockin-redis.txtmeasure-aws/night1-20260812T164952/STATUSmeasure-aws/night1-20260812T164952/winner-cpusnginx/conf.d/default.confscripts/aws-bootstrap.shscripts/e1-tei-cpu-sweep.shscripts/e8-index-sampler.shscripts/night1-run.shsrc/main/java/com/DOCKin/DocKinSpringApplication.javasrc/main/java/com/DOCKin/absence/controller/AbsenceAdminController.javasrc/main/java/com/DOCKin/absence/controller/AbsenceRequestController.javasrc/main/java/com/DOCKin/absence/dto/AbsenceDecisionRequestDto.javasrc/main/java/com/DOCKin/absence/dto/AbsenceRequestCreateRequestDto.javasrc/main/java/com/DOCKin/absence/dto/AbsenceRequestResponseDto.javasrc/main/java/com/DOCKin/absence/event/AbsenceApprovedEvent.javasrc/main/java/com/DOCKin/absence/model/AbsenceRequest.javasrc/main/java/com/DOCKin/absence/model/AbsenceStatus.javasrc/main/java/com/DOCKin/absence/model/AbsenceType.javasrc/main/java/com/DOCKin/absence/repository/AbsenceRequestRepository.javasrc/main/java/com/DOCKin/absence/service/AbsenceRequestService.javasrc/main/java/com/DOCKin/ai/controller/AiController.javasrc/main/java/com/DOCKin/ai/dto/OnlineTranslateDomain.javasrc/main/java/com/DOCKin/ai/model/ChatHistory.javasrc/main/java/com/DOCKin/ai/model/ChatLog.javasrc/main/java/com/DOCKin/ai/model/TranslateLog.javasrc/main/java/com/DOCKin/ai/repository/TranslateRepository.javasrc/main/java/com/DOCKin/ai/service/FastApiService.javasrc/main/java/com/DOCKin/attendance/model/Attendance.javasrc/main/java/com/DOCKin/attendance/model/DayType.javasrc/main/java/com/DOCKin/attendance/model/WorkCalendar.javasrc/main/java/com/DOCKin/attendance/repository/AttendanceRepository.javasrc/main/java/com/DOCKin/attendance/repository/WorkCalendarRepository.javasrc/main/java/com/DOCKin/attendance/service/AbsenceApprovedListener.javasrc/main/java/com/DOCKin/attendance/service/AttendanceBatchService.javasrc/main/java/com/DOCKin/attendance/service/AttendanceService.javasrc/main/java/com/DOCKin/attendance/service/WorkCalendarService.javasrc/main/java/com/DOCKin/checklist/controller/ChecklistAdminController.javasrc/main/java/com/DOCKin/checklist/controller/ChecklistUserController.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistCheckRequestDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistCreateRequestDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistDetailResponseDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistItemRequestDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistItemSimpleResponseDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistItemStatusResponseDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistItemUpdateRequestDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistResponseDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistResultResponseDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistUpdateRequestDto.javasrc/main/java/com/DOCKin/checklist/model/Checklist.javasrc/main/java/com/DOCKin/checklist/model/ChecklistItem.javasrc/main/java/com/DOCKin/checklist/model/ChecklistPhase.javasrc/main/java/com/DOCKin/checklist/model/ChecklistResult.javasrc/main/java/com/DOCKin/checklist/repository/ChecklistItemRepository.javasrc/main/java/com/DOCKin/checklist/repository/ChecklistRepository.javasrc/main/java/com/DOCKin/checklist/repository/ChecklistResultRepository.javasrc/main/java/com/DOCKin/checklist/service/ChecklistService.javasrc/main/java/com/DOCKin/checklist/service/ChecklistStatusService.javasrc/main/java/com/DOCKin/global/config/ClockConfig.javasrc/main/java/com/DOCKin/global/config/RedissonConfig.javasrc/main/java/com/DOCKin/global/config/WebClientConfig.javasrc/main/java/com/DOCKin/global/config/WebConfig.javasrc/main/java/com/DOCKin/global/config/WebSocketConfig.javasrc/main/java/com/DOCKin/global/error/ErrorCode.javasrc/main/java/com/DOCKin/global/error/GlobalExceptionHandler.javasrc/main/java/com/DOCKin/global/file/LogImage.javasrc/main/java/com/DOCKin/global/logging/TraceId.javasrc/main/java/com/DOCKin/global/logging/TraceIdFilter.javasrc/main/java/com/DOCKin/global/security/config/JacksonConfig.javasrc/main/java/com/DOCKin/global/security/config/SecurityConfig.javasrc/main/java/com/DOCKin/global/security/config/SecurityPathConfig.javasrc/main/java/com/DOCKin/global/security/jwt/JwtAuthFilter.javasrc/main/java/com/DOCKin/member/dto/MemberRequestDto.javasrc/main/java/com/DOCKin/member/model/Member.javasrc/main/java/com/DOCKin/member/model/WorkShift.javasrc/main/java/com/DOCKin/member/repository/MemberRepository.javasrc/main/java/com/DOCKin/member/service/MemberService.javasrc/main/java/com/DOCKin/rag/chunking/ChunkingStrategy.javasrc/main/java/com/DOCKin/rag/chunking/FixedSizeChunkingStrategy.javasrc/main/java/com/DOCKin/rag/dto/RetrievalResult.javasrc/main/java/com/DOCKin/rag/dto/RetrievedChunk.javasrc/main/java/com/DOCKin/rag/model/DocumentChunk.javasrc/main/java/com/DOCKin/rag/model/SourceType.javasrc/main/java/com/DOCKin/rag/model/Visibility.javasrc/main/java/com/DOCKin/rag/repository/ChunkVector.javasrc/main/java/com/DOCKin/rag/repository/DocumentChunkRepository.javasrc/main/java/com/DOCKin/rag/repository/NearestChunk.javasrc/main/java/com/DOCKin/rag/service/ChunkIndexWriter.javasrc/main/java/com/DOCKin/rag/service/EmbeddingClient.javasrc/main/java/com/DOCKin/rag/service/IndexingService.javasrc/main/java/com/DOCKin/rag/service/RagChatService.javasrc/main/java/com/DOCKin/rag/service/RetrievalService.javasrc/main/java/com/DOCKin/rag/service/SeedIndexingRunner.javasrc/main/java/com/DOCKin/safetyCourse/model/SafetyCourse.javasrc/main/java/com/DOCKin/safetyCourse/repository/SafetyCourseRepository.javasrc/main/java/com/DOCKin/worklog/controller/WorkLogsController.javasrc/main/java/com/DOCKin/worklog/dto/WorkLogDto.javasrc/main/java/com/DOCKin/worklog/dto/Work_logsDto.javasrc/main/java/com/DOCKin/worklog/model/Comment.javasrc/main/java/com/DOCKin/worklog/model/WorkLog.javasrc/main/java/com/DOCKin/worklog/model/WorkLogImage.javasrc/main/java/com/DOCKin/worklog/model/Work_logs.javasrc/main/java/com/DOCKin/worklog/repository/WorkLogRepository.javasrc/main/java/com/DOCKin/worklog/repository/Work_logsRepository.javasrc/main/java/com/DOCKin/worklog/service/CommentService.javasrc/main/java/com/DOCKin/worklog/service/WorkLogsService.javasrc/main/resources/application-seed.propertiessrc/main/resources/application.propertiessrc/main/resources/data.sqlsrc/main/resources/db/migration/V1__pgvector_and_document_chunks.sqlsrc/main/resources/db/migration/V2__baseline_existing_tables.sqlsrc/main/resources/db/migration/V3__work_log_child_fk_indexes.sqlsrc/main/resources/db/migration/V4__users_child_fk_indexes.sqlsrc/main/resources/db/migration/V5__work_logs_user_id_not_null.sqlsrc/main/resources/db/seed/R__seed_sample.sqlsrc/main/resources/schema.sqlsrc/main/resources/templates/chat_test.htmlsrc/test/java/com/DOCKin/DOCKin_spring/DocKinSpringApplicationTests.javasrc/test/java/com/DOCKin/absence/service/AbsenceRequestServiceTest.javasrc/test/java/com/DOCKin/absence/service/LeaveBalanceConcurrencyTest.javasrc/test/java/com/DOCKin/ai/service/TranslateParallelBenchmarkTest.javasrc/test/java/com/DOCKin/attendance/service/AbsenceApprovedListenerTest.javasrc/test/java/com/DOCKin/attendance/service/AttendanceBatchServiceTest.javasrc/test/java/com/DOCKin/attendance/service/AttendanceServiceTest.javasrc/test/java/com/DOCKin/attendance/service/WorkCalendarServiceTest.javasrc/test/java/com/DOCKin/checklist/service/ChecklistServiceTest.javasrc/test/java/com/DOCKin/checklist/service/ChecklistStatusServiceTest.javasrc/test/java/com/DOCKin/global/config/ActuatorEndpointTest.javasrc/test/java/com/DOCKin/global/config/LocalMigrationDriftTest.javasrc/test/java/com/DOCKin/global/config/SchemaValidationTest.javasrc/test/java/com/DOCKin/global/config/SeedDataTest.javasrc/test/java/com/DOCKin/global/error/GlobalExceptionHandlerTest.javasrc/test/java/com/DOCKin/global/logging/TraceIdFilterTest.javasrc/test/java/com/DOCKin/global/testsupport/ContainerTestSupport.javasrc/test/java/com/DOCKin/global/web/PageableSortDefaultTest.javasrc/test/java/com/DOCKin/member/repository/LockTimeoutVerificationTest.javasrc/test/java/com/DOCKin/rag/chunking/FixedSizeChunkingStrategyTest.javasrc/test/java/com/DOCKin/rag/chunking/SentenceLookbackMeasurementTest.javasrc/test/java/com/DOCKin/rag/repository/FlywayMigrationTest.javasrc/test/java/com/DOCKin/rag/repository/HibernateBatchInsertVerificationTest.javasrc/test/java/com/DOCKin/rag/repository/VectorSearchQueryTest.javasrc/test/java/com/DOCKin/rag/repository/VectorTypeMappingTest.javasrc/test/java/com/DOCKin/rag/service/AnnRecallMeasurementTest.javasrc/test/java/com/DOCKin/rag/service/BruteForceSearchBenchmarkTest.javasrc/test/java/com/DOCKin/rag/service/CorpusSeedGeneratorTest.javasrc/test/java/com/DOCKin/rag/service/CrossLingualRetrievalTest.javasrc/test/java/com/DOCKin/rag/service/EmbeddingClientTest.javasrc/test/java/com/DOCKin/rag/service/RetrievalServiceTest.javasrc/test/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusValidationTest.javasrc/test/java/com/DOCKin/worklog/service/WorkLogListBenchmarkTest.javasrc/test/java/com/DOCKin/worklog/service/WorkLogListQueryCountTest.java
💤 Files with no reviewable changes (9)
- .github/workflows/workflow.yml
- src/main/java/com/DOCKin/ai/model/ChatHistory.java
- src/main/java/com/DOCKin/global/file/LogImage.java
- src/main/java/com/DOCKin/worklog/dto/Work_logsDto.java
- src/main/resources/data.sql
- src/main/java/com/DOCKin/worklog/repository/Work_logsRepository.java
- init.sql
- src/main/java/com/DOCKin/DocKinSpringApplication.java
- src/main/java/com/DOCKin/worklog/model/Work_logs.java
| test: | ||
| name: 빌드 + 테스트 | ||
| runs-on: ubuntu-latest | ||
| timeout-minutes: 20 | ||
|
|
||
| steps: | ||
| - uses: actions/checkout@v7 | ||
|
|
||
| - name: JDK 21 설치 | ||
| uses: actions/setup-java@v5 | ||
| with: | ||
| java-version: '21' | ||
| distribution: 'temurin' | ||
|
|
||
| - name: Gradle 설정 (의존성 캐시 포함) | ||
| uses: gradle/actions/setup-gradle@v6 | ||
|
|
||
| - run: chmod +x gradlew | ||
|
|
||
| # DB와 Redis는 Testcontainers가 띄운다(ContainerTestSupport). ubuntu-latest에는 Docker가 | ||
| # 미리 설치돼 있어 별도 서비스 정의가 필요 없다. | ||
| # | ||
| # 이미지를 먼저 받아두는 것은 실패 원인을 분리하기 위해서다 -- 레지스트리 지연이나 | ||
| # 태그 오타로 테스트가 깨지면 원인이 테스트 안에 있는 것처럼 보인다. | ||
| # 여기서 받아두면 그 경우 이 단계가 먼저 빨간불이 된다. | ||
| # | ||
| # Redis가 목록에 늘어난 이유: @SpringBootTest가 전체 컨텍스트를 요구하는데 RedissonConfig가 | ||
| # 빈 생성 시점에 실제로 접속한다. 2026-08-07부터 CI가 이것 때문에 깨져 있었다(P2-16). | ||
| - name: 테스트용 컨테이너 이미지 준비 | ||
| run: | | ||
| docker pull pgvector/pgvector:pg17 | ||
| docker pull redis:7-alpine | ||
|
|
||
| - name: 테스트 | ||
| run: ./gradlew test --no-daemon | ||
|
|
||
| # 테스트가 실패해도 리포트는 올린다. 실패했을 때가 정작 필요한 순간이다. | ||
| - name: 테스트 리포트 업로드 | ||
| if: always() | ||
| uses: actions/upload-artifact@v7 |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -eu
printf '%s\n' '--- workflow files ---'
git ls-files '.github/workflows/*'
printf '%s\n' '--- ci.yml relevant lines ---'
sed -n '1,110p' .github/workflows/ci.yml
printf '%s\n' '--- workflow permissions and action references ---'
rg -n 'permissions:|contents:|uses:|persist-credentials|gradlew|build.gradle|settings.gradle' .github/workflows . --glob '!build/**' --glob '!node_modules/**' | head -200Repository: DOCKin-project/DOCKin-backend
Length of output: 7457
🏁 Script executed:
#!/bin/bash
set -eu
printf '%s\n' '--- remaining workflow ---'
sed -n '108,180p' .github/workflows/ci.yml
printf '%s\n' '--- action refs with line numbers ---'
nl -ba .github/workflows/ci.yml | sed -n '45,100p'
printf '%s\n' '--- repository build configuration references ---'
rg -n 'pluginManagement|plugins \{|repositories|buildscript|apply plugin|gradlePluginPortal|mavenCentral' settings.gradle* build.gradle* gradle 2>/dev/null | head -160Repository: DOCKin-project/DOCKin-backend
Length of output: 2877
🏁 Script executed:
#!/bin/bash
set -eu
python3 - <<'PY'
import json
import urllib.request
refs = [
("actions", "checkout", "v7"),
("actions", "setup-java", "v5"),
("gradle", "actions", "v6"),
("actions", "upload-artifact", "v7"),
]
for owner, repo, tag in refs:
base = f"https://api.github.com/repos/{owner}/{repo}"
headers = {"Accept": "application/vnd.github+json", "User-Agent": "review-check"}
try:
req = urllib.request.Request(f"{base}/git/ref/tags/{tag}", headers=headers)
ref = json.load(urllib.request.urlopen(req))
obj = ref["object"]
print(f"{owner}/{repo}@{tag}: ref object type={obj['type']} sha={obj['sha']}")
sha = obj["sha"]
if obj["type"] == "tag":
req = urllib.request.Request(f"{base}/git/tags/{sha}", headers=headers)
tag_obj = json.load(urllib.request.urlopen(req))
print(f" dereferenced type={tag_obj['object']['type']} sha={tag_obj['object']['sha']}")
sha = tag_obj["object"]["sha"]
req = urllib.request.Request(f"{base}/commits/{sha}", headers=headers)
commit = json.load(urllib.request.urlopen(req))
print(f" commit endpoint sha={commit['sha']}")
except Exception as e:
print(f"{owner}/{repo}@{tag}: ERROR {e}")
PYRepository: DOCKin-project/DOCKin-backend
Length of output: 822
🏁 Script executed:
#!/bin/bash
set -eu
python3 - <<'PY'
import json, ssl, urllib.request
ctx = ssl._create_unverified_context()
refs = [
("actions", "checkout", "v7"),
("actions", "setup-java", "v5"),
("gradle", "actions", "v6"),
("actions", "upload-artifact", "v7"),
]
for owner, repo, tag in refs:
base = f"https://api.github.com/repos/{owner}/{repo}"
headers = {"Accept": "application/vnd.github+json", "User-Agent": "review-check"}
try:
req = urllib.request.Request(f"{base}/git/ref/tags/{tag}", headers=headers)
ref = json.load(urllib.request.urlopen(req, context=ctx))
obj = ref["object"]
print(f"{owner}/{repo}@{tag}: ref object type={obj['type']} sha={obj['sha']}")
sha = obj["sha"]
if obj["type"] == "tag":
req = urllib.request.Request(f"{base}/git/tags/{sha}", headers=headers)
tag_obj = json.load(urllib.request.urlopen(req, context=ctx))
print(f" dereferenced type={tag_obj['object']['type']} sha={tag_obj['object']['sha']}")
sha = tag_obj["object"]["sha"]
req = urllib.request.Request(f"{base}/commits/{sha}", headers=headers)
commit = json.load(urllib.request.urlopen(req, context=ctx))
print(f" commit endpoint sha={commit['sha']}")
except Exception as e:
print(f"{owner}/{repo}@{tag}: ERROR {e}")
PYRepository: DOCKin-project/DOCKin-backend
Length of output: 426
🏁 Script executed:
#!/bin/bash
set -eu
for spec in \
'actions/checkout v7' \
'actions/setup-java v5' \
'gradle/actions v6' \
'actions/upload-artifact v7'
do
owner=${spec% *}
tag=${spec#* }
printf '%s@%s\n' "$owner" "$tag"
git ls-remote "https://github.com/$owner.git" "refs/tags/$tag" "refs/tags/$tag^{}" || true
doneRepository: DOCKin-project/DOCKin-backend
Length of output: 527
🌐 Web query:
GitHub Actions checkout persist-credentials default true local git config GITHUB_TOKEN documentation
💡 Result:
The persist-credentials input for the GitHub Actions checkout action (actions/checkout) determines whether the repository authentication credentials (such as GITHUB_TOKEN, a PAT, or an SSH key) are configured in the local git config [1][2]. When persist-credentials is set to true (which has historically been the default) [1][3], the action configures the authentication token or SSH key in the local git configuration [1][4]. This process enables subsequent scripts and commands in the job to perform authenticated git operations [1][5]. The action includes a post-job cleanup step to remove these credentials [1][4][5]. Please note that there have been ongoing efforts to change the default value of persist-credentials from true to false to enhance security by reducing the risk of unintentionally exposing the GITHUB_TOKEN secret [6]. You should verify the specific version of the actions/checkout you are using in your workflow file (e.g., v4) to confirm its current default behavior [1][2]. If you wish to disable this behavior, you can explicitly set the input to false: - name: Checkout uses: actions/checkout@v4 with: persist-credentials: false For the most up-to-date information regarding the default settings for your specific version, consult the official actions/checkout repository documentation [1].
Citations:
- 1: https://github.com/actions/checkout
- 2: https://github.com/actions/checkout/tree/v6.0.0
- 3: https://github.com/actions/checkout/blob/72f2cec99f417b1a1c5e2e88945068983b7965f9/action.yml
- 4: https://github.com/actions/checkout/blob/24cb9080177205b6e8c946b17badbe402adc938f/README.md
- 5: https://github.com/actions/checkout/blob/25a956c84d5dd820d28caab9f86b8d183aeeff3d/README.md
- 6: Change the default value of persist-credentials to false actions/checkout#1687
🌐 Web query:
GitHub Actions pin actions to full length commit SHA mutable tags security documentation
💡 Result:
GitHub documentation and security best practices emphasize that pinning GitHub Actions to a full-length (40-character) commit SHA is the only method to ensure an action is used as an immutable release [1][2][3]. Security Risks of Mutable Tags: Version tags (e.g., v1, v4) and branch references (e.g., main) are mutable, meaning they can be moved or overridden at any time [1][4][5]. If an attacker gains access to the repository storing an action, they can retag a version or push malicious code to a branch, causing all workflows that reference that tag to automatically execute the compromised code [6][7][3]. This risk remains even if you trust the original creator of the action [1][2]. Benefits of Full-Length Commit SHA Pinning: - Immutability: A commit SHA points to a specific, unchangeable state of the codebase [4][6][5]. - Protection: Because the SHA is unique to a Git object payload, it mitigates the risk of a malicious actor swapping the underlying code without detection [1][2]. - Governance: GitHub supports policies at the repository and organization levels that can be configured to enforce mandatory SHA pinning, causing any workflow that uses an unpinned action to fail [1][8][3]. Recommended Workflow: To maintain security while still allowing for dependency updates, best practices recommend pinning to the full-length commit SHA while including the human-readable version tag as a trailing comment [6][9][5]. This format allows automated dependency management tools like Dependabot and Renovate to continue tracking and proposing updates based on the tag, while the pipeline itself relies on the immutable SHA [6][5]. Example configuration: - uses: actions/checkout@11bd719 # v4.2.2 This approach ensures that your CI/CD environment is resilient against supply chain compromises involving retagged or malicious upstream actions [9][7].
Citations:
- 1: https://docs.github.com/en/actions/reference/security/secure-use
- 2: https://github.com/github/docs/blob/962a1c8dccb8c0f66548b324e5b921b5e4fbc3d6/content/actions/security-for-github-actions/security-guides/security-hardening-for-github-actions.md
- 3: https://mikael.barbero.tech/blog/post/2026-03-24-stop-trusting-mutable-references/
- 4: https://www.stepsecurity.io/blog/pinning-github-actions-for-enhanced-security-a-complete-guide
- 5: https://nearform.com/insights/why-you-should-pin-your-github-actions-by-commit-hash/
- 6: https://starsling.dev/best-practices/github-actions/pin-action-shas
- 7: https://auditbuffet.com/patterns/ab-002630
- 8: https://github.blog/changelog/2025-08-15-github-actions-policy-now-supports-blocking-and-sha-pinning-actions/
- 9: https://safeguard.sh/resources/blog/github-actions-supply-chain-security
액션을 전체 커밋 SHA로 고정하고 체크아웃 자격 증명 저장을 비활성화하십시오.
actions/checkout@v7, actions/setup-java@v5, gradle/actions/setup-gradle@v6, actions/upload-artifact@v7는 변경 가능한 태그입니다. 각 액션을 버전 주석이 있는 정확한 40자리 커밋 SHA로 고정하십시오.
actions/checkout은 기본적으로 GITHUB_TOKEN을 로컬 Git 설정에 저장합니다. 같은 작업에서 실행하는 PR의 Gradle 빌드 로직이 이 토큰을 읽을 수 있습니다. persist-credentials: false를 설정하고 test 작업에 permissions: contents: read를 명시하십시오.
🧰 Tools
🪛 zizmor (1.29.0)
[warning] 53-53: credential persistence through GitHub Actions artifacts (artipacked): does not set persist-credentials: false
(artipacked)
[warning] 47-110: overly broad permissions (excessive-permissions): default permissions used due to no permissions: block
(excessive-permissions)
[error] 53-53: unpinned action reference (unpinned-uses): action is not pinned to a hash (required by blanket policy)
(unpinned-uses)
[error] 56-56: unpinned action reference (unpinned-uses): action is not pinned to a hash (required by blanket policy)
(unpinned-uses)
[error] 62-62: unpinned action reference (unpinned-uses): action is not pinned to a hash (required by blanket policy)
(unpinned-uses)
[error] 86-86: unpinned action reference (unpinned-uses): action is not pinned to a hash (required by blanket policy)
(unpinned-uses)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.github/workflows/ci.yml around lines 47 - 86, Update the test job’s
actions/checkout, actions/setup-java, gradle/actions/setup-gradle, and
actions/upload-artifact references to 40-character immutable commit SHAs,
retaining version comments for readability. Add test-job permissions with
contents: read, and configure actions/checkout with persist-credentials: false
so the token is not stored locally.
Source: Linters/SAST tools
|
|
||
| ### 동기화/일관성 | ||
|
|
||
| 락은 출근 처리 로직 실행 구간에만 잡고, 트랜잭션 커밋 직후 즉시 해제한다(TTL은 트랜잭션 최대 소요 시간보다 충분히 짧게, 예: 3초). 락 보유 중 커넥션 풀 고갈로 트랜잭션이 지연되는 경우를 대비해 TTL과 DB 커넥션 타임아웃 값의 관계도 함께 점검한다. |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
락 TTL을 보호 구간보다 길게 설정하십시오.
Line 96의 “트랜잭션 최대 소요 시간보다 충분히 짧게”는 반대입니다. TTL이 doClockIn과 DB 커밋 전에 만료되면 다른 요청이 같은 락을 획득할 수 있습니다. TTL은 최대 보호 구간보다 길게 잡거나, 작업 중 lease를 연장하는 watchdog을 사용하십시오.
As per path instructions, “수치가 나오면 그 출처가 실측인지 가정인지 구분되어 있는지 본다.”
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@docs/adr/0001-attendance-clockin-concurrency.md` at line 96, Revise the lock
guidance in the attendance clock-in concurrency ADR so the lock TTL exceeds the
maximum protected doClockIn and database commit duration, or explicitly require
a watchdog to renew the lease while work is active. Replace the incorrect
3-second example and state whether any proposed duration is based on
measurements or an assumption.
Source: Path instructions
| ### 2-3. JVM GC/힙 튜닝 — 근거는 있으나 순서상 보류, 우선순위 3 | ||
|
|
||
| `compose.yaml`에 이미 `JAVA_OPTS=-Xmx400M -Xms256M`, 컨테이너 메모리 512M 제한이 걸려 있다. 즉 "메모리 제약이 있는 환경에서 GC를 어떻게 골랐는가"라는 소재 자체는 근거가 있다. 다만 지금은 GC 로그나 `jstat` 같은 실측 도구가 붙어있지 않은 상태라, 이 상태에서 바로 "GC 튜닝했다"고 쓰면 ADR-0001에서 경계한 "숫자 없이 결론부터 낸" 패턴이 된다. 순서: 먼저 GC 로그/모니터링을 붙이고 실측 → 그 다음 튜닝 여부를 판단. |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
활성 JVM 힙 설정으로 근거를 수정하십시오.
JAVA_OPTS=-Xmx400M은 현재 힙 상한이 아닙니다. Docker ENTRYPOINT가 JAVA_OPTS를 전개하지 않으므로 실제 상한은 -Xmx512M입니다. 이 절의 GC 튜닝 근거와 컨테이너 여유 메모리 설명을 실제 설정 기준으로 갱신하십시오.
As per path instructions, “미검증 설정에 [미검증]을 붙이는 규칙을 쓴다.”
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@docs/adr/0002-performance-improvement-backlog.md` around lines 41 - 43,
Update the “JVM GC/힙 튜닝” section to identify the active heap limit as -Xmx512M,
reflecting the Docker ENTRYPOINT behavior that does not expand JAVA_OPTS. Revise
the supporting explanation of container memory headroom and GC-tuning rationale
to match the effective configuration, and mark any unverified configuration
claims with “[미검증]”.
Source: Path instructions
| if (sequenceCalls > 0 && sequenceCalls <= expectedCalls * 2) { | ||
| System.out.printf(">>> 시퀀스가 %d건씩 묶여 발급된다. IDENTITY 시절이라면 %,d회였을 자리다.%n", | ||
| ALLOCATION_SIZE, TOTAL); | ||
| } else if (sequenceCalls >= TOTAL) { | ||
| System.out.println(">>> 시퀀스가 건별로 호출되고 있다. allocationSize가 적용되지 않았다."); | ||
| } else { | ||
| System.out.println(">>> pg_stat_statements에서 관측하지 못했다(확장 미설치 등)."); | ||
| } | ||
| System.out.println(); | ||
|
|
||
| assertEquals(ALLOCATION_SIZE, incrementBy, | ||
| "시퀀스 increment_by가 allocationSize와 다르다 - 설정이 DB에 반영되지 않았다"); | ||
| if (sequenceCalls > 0) { | ||
| assertTrue(sequenceCalls < TOTAL / 2, | ||
| "시퀀스가 " + sequenceCalls + "회 호출됐다 - allocationSize가 적용되지 않았다"); | ||
| } | ||
| } catch (SQLException e) { | ||
| Assumptions.abort("테스트 컨테이너 접속 실패로 검증을 건너뜁니다: " + e.getMessage()); | ||
| } | ||
| } | ||
|
|
||
| private void resetStatements(Connection conn) { | ||
| try (Statement st = conn.createStatement()) { | ||
| st.execute("SELECT pg_stat_statements_reset()"); | ||
| } catch (SQLException ignored) { | ||
| // 확장이 없으면 시퀀스 관측을 포기하고 시간만 본다. | ||
| } |
There was a problem hiding this comment.
🚀 Performance & Scalability | 🟠 Major | ⚡ Quick win
관측 실패와 과도한 시퀀스 호출을 통과로 처리하지 마십시오.
pg_stat_statements 초기화 실패는 무시됩니다. 이후 sequenceCalls == 0이면 핵심 검사를 건너뜁니다.
또한 1,000회 같은 호출 수는 기대값 약 200회보다 크게 나쁘지만 현재 sequenceCalls < 5,000 조건을 통과합니다. 관측 확장 실패는 테스트 실패로 처리하십시오. sequenceCalls > 0을 요구하고, 상한은 출력 조건과 같이 expectedCalls * 2 수준으로 제한하십시오.
As per path instructions, 선언된 설정이나 애노테이션이 실제로 효과가 있는지 확인해야 합니다.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In
`@src/test/java/com/DOCKin/rag/repository/HibernateBatchInsertVerificationTest.java`
around lines 124 - 150, Update HibernateBatchInsertVerificationTest so
resetStatements(Connection) propagates pg_stat_statements reset failures instead
of ignoring SQLException, causing the test to fail when observation is
unavailable. In the verification assertions, require sequenceCalls > 0 and
enforce the same expectedCalls * 2 upper bound used by the diagnostic output,
replacing the permissive TOTAL / 2 threshold.
Source: Path instructions
| try { | ||
| var violations = validator.validate(dto); | ||
| System.out.println("결과: 예외 없이 통과했다."); | ||
| System.out.println("위반 건수: " + violations.size()); | ||
| violations.forEach(v -> System.out.println(" - " + v.getPropertyPath() + ": " + v.getMessage())); | ||
| System.out.println(); | ||
| System.out.println("→ 검증이 걸리지 않는다. courseId가 null이어도 서비스까지 내려간다."); | ||
| } catch (RuntimeException e) { | ||
| System.out.println("결과: 예외가 났다."); | ||
| System.out.println(" 예외 : " + e.getClass().getName()); | ||
| System.out.println(" 메시지: " + e.getMessage()); | ||
| System.out.println(); | ||
| System.out.println("→ @Valid가 붙은 컨트롤러 경로는 요청마다 이 예외를 맞는다."); | ||
| System.out.println(" GlobalExceptionHandler의 캐치올에 걸리면 500이 된다."); | ||
| } |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
fd -a 'SafetyWatchStatusRequestDto.java' src | while IFS= read -r file; do
rg -n -C 3 '`@NotBlank`|`@NotNull`|courseId' "$file"
done
rg -n -C 3 'validator\.validate|assertThrows|assertEquals|assertTrue' \
src/test/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusValidationTest.javaRepository: DOCKin-project/DOCKin-backend
Length of output: 822
🏁 Script executed:
#!/bin/bash
set -euo pipefail
printf '%s\n' '--- DTO ---'
fd -a 'SafetyWatchStatusRequestDto.java' . | while IFS= read -r file; do
cat -n "$file"
done
printf '%s\n' '--- validation test ---'
cat -n src/test/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusValidationTest.java
printf '%s\n' '--- controller and DTO usages ---'
rg -n -C 5 'SafetyWatchStatusRequestDto|completeCourse|`@Valid`' src/main/java
printf '%s\n' '--- validation dependencies ---'
rg -n -C 2 'validation|hibernate-validator|jakarta.validation|javax.validation' \
build.gradle* pom.xml gradle.properties settings.gradle* 2>/dev/null || trueRepository: DOCKin-project/DOCKin-backend
Length of output: 45322
🏁 Script executed:
#!/bin/bash
set -euo pipefail
python3 - <<'PY'
from pathlib import Path
import re
dto = Path("src/main/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusRequestDto.java").read_text()
test = Path("src/test/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusValidationTest.java").read_text()
controller = Path("src/main/java/com/DOCKin/safetyCourse/controller/SafetyUserController.java").read_text()
field = re.search(
r'(?P<annotations>(?:\s*@\w+(?:\([^)]*\))?\s*)*)'
r'private\s+(?P<type>\w+)\s+(?P<name>courseId)\s*;',
dto,
)
assert field, "courseId declaration not found"
annotations = field.group("annotations")
print(f"courseId type: {field.group('type')}")
print(f"courseId annotations: {re.findall(r'@\w+', annotations)}")
print(f"test assertion calls: {len(re.findall(r'\bassert(?:Equals|True|False|Null|NotNull|Throws|All)\b', test))}")
print(f"test catches RuntimeException: {bool(re.search(r'catch\s*\(\s*RuntimeException\b', test))}")
print(f"completeCourse uses `@Valid`: {bool(re.search(r'`@Valid`\s+@RequestBody\s+SafetyWatchStatusRequestDto', controller))}")
assert field.group("type") == "Integer"
assert "`@NotBlank`" in annotations
assert not re.search(r'\bassert(?:Equals|True|False|Null|NotNull|Throws|All)\b', test)
assert re.search(r'catch\s*\(\s*RuntimeException\b', test)
assert re.search(r'`@Valid`\s+@RequestBody\s+SafetyWatchStatusRequestDto', controller)
PYRepository: DOCKin-project/DOCKin-backend
Length of output: 400
🏁 Script executed:
#!/bin/bash
set -euo pipefail
python3 - <<'PY'
from pathlib import Path
import re
dto = Path("src/main/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusRequestDto.java").read_text()
test = Path("src/test/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusValidationTest.java").read_text()
controller = Path("src/main/java/com/DOCKin/safetyCourse/controller/SafetyUserController.java").read_text()
field = re.search(
r'(?P<annotations>(?:\s*@\w+(?:\([^)]*\))?\s*)*)'
r'private\s+(?P<type>\w+)\s+(?P<name>courseId)\s*;',
dto,
)
assert field, "courseId declaration not found"
annotations = field.group("annotations")
annotation_names = re.findall(r'@\w+', annotations)
assertions = re.findall(
r'\bassert(?:Equals|True|False|Null|NotNull|Throws|All)\b',
test,
)
print(f"courseId type: {field.group('type')}")
print(f"courseId annotations: {annotation_names}")
print(f"test assertion calls: {len(assertions)}")
print(f"test catches RuntimeException: {bool(re.search(r'catch\s*\(\s*RuntimeException\b', test))}")
print(
"completeCourse uses `@Valid`: "
f"{bool(re.search(r'`@Valid`\s+@RequestBody\s+SafetyWatchStatusRequestDto', controller))}"
)
assert field.group("type") == "Integer"
assert "`@NotBlank`" in annotation_names
assert not assertions
assert re.search(r'catch\s*\(\s*RuntimeException\b', test)
assert re.search(r'`@Valid`\s+@RequestBody\s+SafetyWatchStatusRequestDto', controller)
PYRepository: DOCKin-project/DOCKin-backend
Length of output: 468
🏁 Script executed:
#!/bin/bash
set -euo pipefail
python3 - <<'PY'
from pathlib import Path
import re
dto = Path("src/main/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusRequestDto.java").read_text()
test = Path("src/test/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusValidationTest.java").read_text()
controller = Path("src/main/java/com/DOCKin/safetyCourse/controller/SafetyUserController.java").read_text()
field = re.search(
r'(?P<annotations>(?:\s*@\w+(?:\([^)]*\))?\s*)*)'
r'private\s+(?P<type>\w+)\s+(?P<name>courseId)\s*;',
dto,
)
assert field, "courseId declaration not found"
annotations = field.group("annotations")
annotation_names = re.findall(r'@\w+', annotations)
assertions = re.findall(
r'\bassert(?:Equals|True|False|Null|NotNull|Throws|All)\b',
test,
)
catches_runtime_exception = bool(
re.search(r'catch\s*\(\s*RuntimeException\b', test)
)
uses_valid = bool(
re.search(r'`@Valid`\s+@RequestBody\s+SafetyWatchStatusRequestDto', controller)
)
print("courseId type:", field.group("type"))
print("courseId annotations:", annotation_names)
print("test assertion calls:", len(assertions))
print("test catches RuntimeException:", catches_runtime_exception)
print("completeCourse uses `@Valid`:", uses_valid)
assert field.group("type") == "Integer"
assert "`@NotBlank`" in annotation_names
assert not assertions
assert catches_runtime_exception
assert uses_valid
PYRepository: DOCKin-project/DOCKin-backend
Length of output: 330
courseId의 제약 조건과 검증 테스트를 수정하십시오.
SafetyWatchStatusRequestDto.courseId는 Integer인데 @NotBlank가 선언되어 있습니다. @NotBlank는 CharSequence 대상이므로 @Valid가 적용된 completeCourse 검증에서 타입 불일치 예외가 발생합니다. 현재 테스트는 정상 반환과 RuntimeException을 모두 출력만 하며 assertion 없이 종료합니다. @NotNull로 변경하고, null 값의 violation과 유효한 정수의 무위반을 assertion으로 검증하십시오.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In
`@src/test/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusValidationTest.java`
around lines 41 - 55, Update SafetyWatchStatusRequestDto.courseId to use
`@NotNull` instead of `@NotBlank`, matching its Integer type and preventing
validation type errors. In SafetyWatchStatusValidationTest, replace print-only
handling with assertions that null courseId produces a violation while a valid
integer produces none; keep the test failing on unexpected RuntimeException.
| } finally { | ||
| cleanUp(conn); | ||
| } | ||
| } catch (SQLException e) { | ||
| Assumptions.abort("PostgreSQL(localhost:5432)에 접속할 수 없어 건너뜁니다: " + e.getMessage()); | ||
| } | ||
| } |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
측정 중 발생한 SQLException이 "접속 불가"로 삼켜져 테스트가 건너뜀으로 보고된다.
189-191행의 catch는 try-with-resources 블록 전체를 감싼다. getConnection의 실패만 잡는 것이 아니다. load, measureAll, createBenchIndexes, cleanUp 중 어디서 나는 SQLException도 여기로 온다.
그러면 Assumptions.abort가 호출되어 테스트가 건너뜀으로 끝난다. 벤치마크가 실제로 깨져도 실행 결과는 "PostgreSQL에 접속할 수 없어 건너뜁니다"가 된다. 원인 메시지가 사실과 다르고, 실패가 성공적 skip으로 위장된다.
187행의 cleanUp이 실패하는 경우가 특히 문제다. 100만 행의 벤치 데이터가 남았는데도 테스트는 조용히 건너뜀으로 끝난다. 42-44행의 주석이 경계한 상황("벤치 행이 남은 채 앱이 뜨면 cron이 그것들을 임베딩하려 든다")이 바로 그것이다.
접속 실패만 별도로 판정해 주십시오.
🐛 제안 수정 — 접속 단계와 측정 단계를 분리한다
- try (Connection conn = DriverManager.getConnection(URL, "root", password)) {
+ Connection opened;
+ try {
+ opened = DriverManager.getConnection(URL, "root", password);
+ } catch (SQLException e) {
+ Assumptions.abort("PostgreSQL(localhost:5432)에 접속할 수 없어 건너뜁니다: " + e.getMessage());
+ return;
+ }
+ // 이 아래의 SQLException은 측정 실패다. 건너뜀으로 바꾸지 않고 그대로 던진다.
+ try (Connection conn = opened) {
try {
cleanUp(conn);
// ... 이하 동일 ...
} finally {
cleanUp(conn);
}
- } catch (SQLException e) {
- Assumptions.abort("PostgreSQL(localhost:5432)에 접속할 수 없어 건너뜁니다: " + e.getMessage());
}📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| } finally { | |
| cleanUp(conn); | |
| } | |
| } catch (SQLException e) { | |
| Assumptions.abort("PostgreSQL(localhost:5432)에 접속할 수 없어 건너뜁니다: " + e.getMessage()); | |
| } | |
| } | |
| Connection opened; | |
| try { | |
| opened = DriverManager.getConnection(URL, "root", password); | |
| } catch (SQLException e) { | |
| Assumptions.abort("PostgreSQL(localhost:5432)에 접속할 수 없어 건너뜁니다: " + e.getMessage()); | |
| return; | |
| } | |
| // 이 아래의 SQLException은 측정 실패다. 건너뜀으로 바꾸지 않고 그대로 던진다. | |
| try (Connection conn = opened) { | |
| try { | |
| cleanUp(conn); | |
| // ... 이하 동일 ... | |
| } finally { | |
| cleanUp(conn); | |
| } | |
| } |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/test/java/com/DOCKin/worklog/service/WorkLogListBenchmarkTest.java`
around lines 186 - 192, 분석벤치마크 테스트의 전체 실행을 감싸는 SQLException catch를 제거하고,
getConnection 단계에서 발생한 접속 실패만 Assumptions.abort로 처리하도록 연결 획득과 측정·정리 단계(load,
measureAll, createBenchIndexes, cleanUp)를 분리하십시오. 측정 또는 cleanUp에서 발생한
SQLException은 건너뛰지 말고 테스트 실패로 전파되게 하며, 기존 정리 동작은 유지하십시오.
| private void cleanUp(Connection conn) throws SQLException { | ||
| dropBenchIndexes(conn); | ||
| try (Statement st = conn.createStatement()) { | ||
| st.executeUpdate(""" | ||
| DELETE FROM work_log_images | ||
| WHERE work_log_id IN (SELECT log_id FROM work_logs WHERE title LIKE '%s%%') | ||
| """.formatted(MARKER)); | ||
| st.executeUpdate("DELETE FROM work_logs WHERE title LIKE '%s%%'".formatted(MARKER)); | ||
| st.executeUpdate("DELETE FROM users WHERE user_id LIKE '%s%%'".formatted(BENCH_USER_PREFIX)); | ||
|
|
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
LIKE 'bch%'는 이 테스트가 만들지 않은 사용자까지 지운다.
576행은 진짜 users 테이블에서 user_id LIKE 'bch%'인 행을 모두 지운다. 이 테스트가 만드는 ID는 bch0000 형식으로 접두사 3자 + 숫자 4자다(369행). 그러나 삭제 패턴은 그보다 넓다. bch로 시작하는 실제 사용자 ID가 개발 DB에 있으면 함께 지워진다.
573행과 575행의 MARKER는 [벤치] 로 충돌 가능성이 낮다. bch는 다르다. 짧고, 사람이 쓸 수 있는 문자열이다.
이 테스트는 복제본이 아니라 운영 스키마의 진짜 테이블을 대상으로 한다(33-36행). 삭제 범위를 이 테스트가 만드는 형식과 정확히 일치시켜 주십시오.
🛡️ 제안 수정 — 삭제 범위를 생성 형식으로 좁힌다
- st.executeUpdate("DELETE FROM users WHERE user_id LIKE '%s%%'".formatted(BENCH_USER_PREFIX));
+ // 생성 형식(bch + 숫자 4자)과 정확히 일치하는 행만 지운다.
+ // 'bch%'로 넓히면 이 테스트가 만들지 않은 사용자까지 지운다.
+ st.executeUpdate("DELETE FROM users WHERE user_id ~ '^%s[0-9]{4}$'"
+ .formatted(BENCH_USER_PREFIX));🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/test/java/com/DOCKin/worklog/service/WorkLogListBenchmarkTest.java`
around lines 568 - 577, Update the user deletion statement in
cleanUp(Connection) to match exactly the benchmark IDs generated by the test:
the bch prefix followed by four digits, rather than any user_id beginning with
bch. Keep the existing cleanup of benchmark-created users while excluding
unrelated user records.
| Measurement all10 = measure("전체 목록", 10, | ||
| () -> workLogsService.readWorklog(user(0), PageRequest.of(0, 10))); | ||
| Measurement all20 = measure("전체 목록", 20, | ||
| () -> workLogsService.readWorklog(user(0), PageRequest.of(0, 20))); | ||
| Measurement other = measure("타인 목록", 20, | ||
| () -> workLogsService.readOtherWorklog(otherViewer(), otherTarget(), PageRequest.of(0, 20))); | ||
| Measurement search = measure("키워드 검색", 20, | ||
| () -> workLogsService.searchByKeyword(KEYWORD, PageRequest.of(0, 20))); | ||
|
|
||
| print(all10, all20, other, search); | ||
|
|
||
| assertEquals(20, all20.rows(), "표본이 한 페이지를 채우지 못하면 행당 비용을 볼 수 없다"); | ||
| assertEquals(20, other.rows(), "타인 목록 표본이 한 페이지를 채우지 못했다"); | ||
| assertEquals(20, search.rows(), "키워드 검색 표본이 한 페이지를 채우지 못했다"); | ||
|
|
||
| assertEquals(4 + 10, all10.queries(), explain("전체 목록(10)", all10)); | ||
| assertEquals(4 + 20, all20.queries(), explain("전체 목록(20)", all20)); | ||
| assertEquals(4 + 20, other.queries(), explain("타인 목록(20)", other)); | ||
| assertEquals(2 + 20, search.queries(), explain("키워드 검색(20)", search)); |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
측정 호출과 전제 검사의 정렬 기준을 일치시키십시오.
Line 160과 Line 162는 정렬 없는 PageRequest를 사용합니다. Line 198부터 Line 204는 ORDER BY log_id로 다른 첫 페이지를 검사합니다.
정렬 없는 SQL 결과 순서는 보장되지 않습니다. 따라서 현재 전제 검사는 실제 측정 페이지의 작성자와 장비가 모두 다른지 보장하지 못합니다.
측정 호출에 Sort.by("logId")를 지정하십시오.
수정 예시
-() -> workLogsService.readWorklog(user(0), PageRequest.of(0, 10)));
+() -> workLogsService.readWorklog(user(0),
+ PageRequest.of(0, 10, Sort.by("logId"))));
-() -> workLogsService.readWorklog(user(0), PageRequest.of(0, 20)));
+() -> workLogsService.readWorklog(user(0),
+ PageRequest.of(0, 20, Sort.by("logId"))));Also applies to: 195-211
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/test/java/com/DOCKin/worklog/service/WorkLogListQueryCountTest.java`
around lines 159 - 177, Update the paginated measurement calls in the test,
including the all10, all20, other, and search cases, to use PageRequest.of with
Sort.by("logId"). Keep the measurement assertions unchanged and align these
requests with the ORDER BY log_id assumptions used by the later precondition
checks.
There was a problem hiding this comment.
Actionable comments posted: 40
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (2)
src/main/java/com/DOCKin/worklog/service/WorkLogsService.java (1)
40-69: 🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift원격 작업을 트랜잭션 밖으로 분리하십시오.
createWorklog와createSttWorklog는 DB 조회 후 같은@Transactional범위에서uploadImage와processStt(...).block()을 실행합니다. 원격 작업이 지연되면 DB 트랜잭션과 커넥션이 원격 응답까지 유지됩니다. STT 클라이언트의 제한 시간 설정은 이 문제를 해결하지 못합니다.검증과 저장을 짧은 DB 트랜잭션으로 분리하고, STT 처리와 S3 업로드를 트랜잭션 밖에서 실행하십시오. 저장에 실패하면 생성한 S3 객체를 삭제하십시오.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/main/java/com/DOCKin/worklog/service/WorkLogsService.java` around lines 40 - 69, Refactor createWorklog and createSttWorklog so validation and WorkLog persistence occur in a short transaction that ends before invoking uploadImage or processStt(...).block(); perform STT processing and S3 uploads outside the transactional methods, then persist the resulting data in a separate transaction, and delete any S3 objects created during the operation if saving fails.Source: Path instructions
src/main/java/com/DOCKin/ai/service/FastApiService.java (1)
124-153: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win외부 HTTP 호출을 트랜잭션 밖으로 이동하십시오.
saveTranslateLog는@Transactional상태에서Mono.zip(...).block()으로 FastAPI 응답을 기다립니다. FastAPI가 지연되면 DB 트랜잭션과 커넥션이 응답 시간 전체에 걸쳐 점유됩니다. 커넥션 풀이 소진되면 다른 요청도 중단됩니다.원본 작업일지 조회와 FastAPI 호출을 트랜잭션 밖에서 수행하십시오. 번역 행 upsert만 짧은 쓰기 트랜잭션으로 처리하십시오. 기존 동시 재번역 원자성 요구사항도 유지하십시오.
As per path instructions, “트랜잭션 안에서 외부 HTTP를 호출하면 반드시 지적한다.”
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/main/java/com/DOCKin/ai/service/FastApiService.java` around lines 124 - 153, saveTranslateLog에서 `@Transactional` 범위를 제거해 WorkLog 조회와 FastAPI 호출(Mono.zip(...).block())이 트랜잭션 밖에서 수행되도록 분리하십시오. 번역 응답을 받은 뒤 번역 행 upsert만 별도의 짧은 쓰기 트랜잭션 메서드로 위임하고, 동시 재번역 시 원자성이 유지되도록 기존 upsert 처리와 트랜잭션 경계를 보존하십시오.Source: Path instructions
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@build.gradle`:
- Around line 106-107: Update the S3 dependency in the build configuration from
the legacy spring-cloud-starter-aws 2.2.6.RELEASE to the Spring Cloud AWS
starter and version compatible with Spring Boot 4.0.1, following the official
compatibility matrix. Then adapt the S3 client initialization and related APIs
to the selected starter’s current API.
In `@compose.gc.yaml`:
- Around line 29-32: Update the JAVA_OPTS assignment in the GC overlay to retain
the Dockerfile’s -Xmx400M heap limit alongside the existing GC logging options.
Verify the container ENTRYPOINT consumes this environment variable so the
measurement run preserves the same JVM memory conditions as normal execution.
In `@docs/2026-07-04-checklist-domain-work-summary.md`:
- Line 17: 문서의 측정 상태 표기를 근거와 함께 일관되게 보완하십시오.
docs/2026-07-04-checklist-domain-work-summary.md 17-17의 테스트 수치에 실행 시점과 실행 명령 또는
CI 산출물 경로를 추가하고, docs/AWS-MEASUREMENT-PLAN.md 81-81의 780MB·390MB 추정에는 산정 근거와 [실측
필요]를 표시하십시오. docs/AWS-MEASUREMENT-RESULTS.md 90-93의 40MB·50MB·300MB 추정에는 원본 측정
경로를 연결하거나 [실측 필요]를 추가하며, 가정에는 [실측 필요], 미검증 설정에는 [미검증] 표기를 적용하십시오.
In `@docs/adr/0006-rag-vector-search.md`:
- Line 485: Update the language-performance statement near the ADR’s comparison
to mark the unsupported claim about real-world language usage with “[실측 필요]” or
add evidence, consistent with ADR-0003’s lack of traffic data. Also mark the
unmeasured “2~3배” chunk multiplier near the referenced comparison with “[실측
필요]”, following the document’s labeling rule for assumptions and unverified
settings.
In `@docs/adr/0007-corpus-retention.md`:
- Line 14: ADR-0007의 일일·연간 작업일지 수치를 동일한 산식으로 정렬하십시오. 4,000건/일을 유지하면 연간 약 146만
건으로 수정하고, 100만 건을 유지하려면 약 250 운영일이라는 가정을 명시하십시오. 해당 수치가 실측인지 가정인지 출처를 구분해 표기하고,
4절의 청크·저장 용량 외삽 및 파티션 계획에도 일관된 값을 반영하십시오.
- Around line 23-27: ADR의 “만료 구간 검색” 항목을 확정 결정이 아닌 미결정 또는 [미검증] 상태로 변경하고, 키워드 경로
폴백을 확정한 표현을 제거하십시오. 7절의 관련 설명(114-131 구간)도 현재 키워드 경로가 삭제된 document_chunks를 조회하며
원본 검색 경로와 권한 선필터가 미결정임을 반영하도록 수정하고, 원본 검색 경로 구현 후에만 폴백 정책을 확정한다고 기록하십시오.
In `@docs/PORTFOLIO-ROADMAP.md`:
- Around line 480-488: Update docs/PORTFOLIO-ROADMAP.md lines 480-488 so the
FastAPI timeout statement reflects the completed N1/N3 work or is explicitly
labeled as a historical diagnostic snapshot; also update lines 649-657 so the
WebSocket/Nginx statement is likewise current or clearly historical. Preserve
the completed-status record around lines 387-390 and remove any presentation of
these resolved issues as active defects.
In `@scripts/night1-run.sh`:
- Around line 218-227: Replace hardcoded dockin-app-1 container references with
IDs resolved from the compose service: in scripts/night1-run.sh lines 218-227
and 124-128, use dc ps -qa "$SVC_APP"; in scripts/aws-bootstrap.sh lines
166-171, use docker compose $COMPOSE_FILES ps -qa dockin-app for both health
checks. Preserve the existing docker logs and health-status behavior while
ensuring all three sites work with compose-generated container names.
In `@src/main/java/com/DOCKin/absence/service/AbsenceRequestService.java`:
- Around line 56-68: createRequest에서 s3PresignedService.uploadImage 호출을
`@Transactional` 범위 밖으로 분리하고, DB 저장 실패 시 이미 업로드된 객체를 삭제하는 보상 처리를 추가하십시오. 실제 S3
클라이언트의 연결 및 읽기 타임아웃도 설정해 외부 호출이 무기한 대기하지 않도록 하십시오.
- Around line 47-54: Make the state transition for a request identified by
requirePendingRequest atomic so concurrent transactions cannot both process the
same PENDING request. Acquire a database lock by requestId while loading the
request, then recheck its status under that lock before returning it;
alternatively use a single conditional PENDING-to-processed update and verify
its affected-row count. Ensure this request lock is obtained before any
vacation-balance lock used by the approval flows.
In `@src/main/java/com/DOCKin/attendance/service/AbsenceApprovedListener.java`:
- Around line 55-65: 승인 처리와 자정 배치가 동일한 날짜를 원자적으로 조정하도록 수정하십시오.
src/main/java/com/DOCKin/attendance/service/AbsenceApprovedListener.java
55-65에서는 기존 기록을 모두 건너뛰지 말고 자동 생성된 ABSENT만 승인 상태(VACATION 또는 SICK)로 갱신하며 실제 출퇴근
기록은 유지하십시오.
src/main/java/com/DOCKin/attendance/service/AttendanceBatchService.java 85-95에서도
동일한 (user_id, work_date)에 대해 원자적 충돌 처리와 상태 우선순위를 적용해 승인된 휴가가 ABSENT로 덮어써지지 않게
하십시오.
In `@src/main/java/com/DOCKin/attendance/service/AttendanceService.java`:
- Around line 122-124: Update the attendance lookup in AttendanceService to
retrieve the member’s most recent open attendance record (without a checkout
time) instead of querying by workDate=today. Ensure the repository query selects
at most one record for the member, preserving ATTENDANCE_NOT_CHECKED_IN when no
open record exists.
- Around line 126-132: AttendanceService의 퇴근 처리에서 조회 후 clockOutTime을 설정하는 비원자적
흐름을 수정하십시오. clockOutTime IS NULL 조건을 포함한 단일 UPDATE를 사용하고 영향받은 행이 0이면
ATTENDANCE_ALREADY_CHECKED_OUT을 발생시키도록 하거나, 기존 조회에 행 잠금을 적용해 동시 요청을 직렬화하십시오.
In `@src/main/java/com/DOCKin/checklist/controller/ChecklistAdminController.java`:
- Around line 40-44: 관리자 체크리스트 상세 조회에 ADMIN 권한 검증을 추가하십시오.
ChecklistAdminController의 getChecklist와 ChecklistService의 getChecklist 흐름에서 기존
requireAdmin 권한 검증 방식을 재사용하여 일반 인증 사용자의 접근을 차단하십시오.
In `@src/main/java/com/DOCKin/checklist/dto/ChecklistUpdateRequestDto.java`:
- Around line 17-18: Update the title field in ChecklistUpdateRequestDto to
reject empty or whitespace-only values while continuing to allow null for
partial updates. Use the project’s existing validation approach so `@Valid`
enforces this contract before ChecklistService.updateChecklist() persists the
value.
In `@src/main/java/com/DOCKin/checklist/service/ChecklistService.java`:
- Around line 65-67: Update ChecklistService.createChecklist to catch
DataIntegrityViolationException from the checklist persistence operation,
including saveAndFlush or the applicable transaction boundary, and convert
duplicate equipment_id/phase violations into
BusinessException(ErrorCode.CHECKLIST_ALREADY_EXISTS). Preserve the existing
pre-check while ensuring concurrent duplicate requests no longer surface as 500
responses.
In `@src/main/java/com/DOCKin/checklist/service/ChecklistStatusService.java`:
- Around line 49-62: Update the query behind
checklistResultRepository.findLatestResultsByChecklistId to eagerly load the
associated member using JOIN FETCH or an equivalent `@EntityGraph`. Ensure
latest.getMember().getUserId() in ChecklistStatusService does not trigger an
additional SELECT for each result.
In `@src/main/java/com/DOCKin/global/config/WebClientConfig.java`:
- Around line 22-35: Refactor ChunkIndexWriter#writePage so the embedding HTTP
call runs outside any database transaction, then open a short transaction only
for persisting the embedding results. Preserve the existing write behavior while
ensuring no transactional method or transaction-scoped operation surrounds the
external embedding request.
In `@src/main/java/com/DOCKin/global/logging/TraceIdFilter.java`:
- Around line 50-55: Update TraceIdFilter so the response header is refreshed
with the final MDC trace ID after TraceId.override() processing and before the
response is committed, ensuring it matches logs and chat_history.trace_id. Add a
regression test covering an override and asserting the response X-Trace-Id
equals the final trace ID.
In `@src/main/java/com/DOCKin/global/security/config/SecurityConfig.java`:
- Around line 33-37: SecurityConfig.setupSecurityContext()에서
MODE_INHERITABLETHREADLOCAL 설정을 제거하고 SecurityContextHolder의 기본 전략을 유지하십시오. 인증
컨텍스트가 필요한 `@Async` 실행 경로는 DelegatingSecurityContextAsyncTaskExecutor 계열 executor로
구성해 작업 제출 시점에 컨텍스트를 전파하고 실행 후 정리하도록 변경하십시오.
In `@src/main/java/com/DOCKin/member/model/Member.java`:
- Around line 66-70: Update Member.useLeaveDays to reject days values less than
or equal to zero before checking remainingLeaveDays, using the appropriate
input-validation BusinessException/ErrorCode; preserve the existing
insufficient-balance handling for positive values that exceed
remainingLeaveDays.
- Around line 45-52: 기존 users 테이블을 갱신하는 새 버전 마이그레이션을 추가하십시오. 엔티티의 workShift와
remainingLeaveDays에 대응하는 컬럼을 추가하고, 기존 행의 remaining_leave_days를 15로 백필하십시오.
work_shift의 NOT NULL 적용 여부는 현재 nullable 정책과 일치하도록 별도로 결정하며, 새 컬럼과 백필이
ddl-auto=validate 스키마와 일치하도록 구성하십시오.
In `@src/main/java/com/DOCKin/member/repository/MemberRepository.java`:
- Around line 47-53: Update the Javadoc surrounding
MemberRepository.findByUserIdForUpdate to describe PostgreSQL locking and the
actual lock_timeout=5s setting instead of MySQL/InnoDB and compose.yaml
innodb-lock-wait-timeout terminology, while preserving the explanation that the
lock read, balance check, and update share approveRequest’s transaction.
In `@src/main/java/com/DOCKin/rag/model/DocumentChunk.java`:
- Around line 166-182: Validate the embedding dimension in both the
DocumentChunk constructor and reindex path before applying state changes,
requiring the actual embedding length to equal embeddingDim. Reject mismatched
values, including cases where either value is missing when validation is
required, while preserving valid updates.
In `@src/main/java/com/DOCKin/rag/service/RetrievalService.java`:
- Around line 70-72: 고정된 CANDIDATE_MULTIPLIER만으로는 원본별 제한 후에도 topK를 보장할 수 없으므로
RetrievalService의 후보 조회를 원본별 SQL 랭킹 제한 방식 또는 topK가 충족될 때까지 후보를 확장하는 방식으로 변경하십시오.
src/main/java/com/DOCKin/rag/service/RetrievalService.java 70-72의 후보 수 로직을 수정하고,
src/test/java/com/DOCKin/rag/service/RetrievalServiceTest.java 127-140에는 한 원본이
후보 창을 독점해도 최종 결과가 topK를 채우는 결과 기반 회귀 테스트를 추가하십시오.
- Around line 134-140: Remove `@Transactional` from RetrievalService.search so
embeddingClient.embedQuery runs outside any transaction. Generate the query
embedding first, then delegate repository retrieval to a separate Spring bean
with a transactional read-only method; do not use self-invocation, so the
transaction proxy is applied to the database-only operation.
In `@src/main/java/com/DOCKin/rag/service/SeedIndexingRunner.java`:
- Around line 29-33: Update the Javadoc condition description in
SeedIndexingRunner to name the seed profile instead of local, matching the
`@Profile`("seed") condition while leaving the indexing property description
unchanged.
In `@src/main/java/com/DOCKin/worklog/model/Comment.java`:
- Around line 25-27: Rename the Comment entity association field from logId to
workLog because it stores a WorkLog entity, then update its getter, builder
usage, and any Spring Data derived queries accordingly. In
src/main/java/com/DOCKin/worklog/model/Comment.java lines 25-27, rename the
field; in src/main/java/com/DOCKin/worklog/service/CommentService.java lines
43-44, change the builder call to use workLog(workLog) and update related
accessors or query property names.
In `@src/main/resources/db/migration/V4__users_child_fk_indexes.sql`:
- Around line 25-52: work_logs의 대용량 인덱스 생성이 쓰기 작업을 차단하지 않도록 처리하십시오. 기존 V4
마이그레이션은 수정하지 말고, 트랜잭션 없이 실행되는 새 버전 마이그레이션에서 idx_work_logs_user를 CREATE INDEX
CONCURRENTLY IF NOT EXISTS로 생성하십시오. 기존 데이터베이스에도 안전하게 적용되도록 멱등성을 유지하십시오.
In `@src/main/resources/db/seed/R__seed_sample.sql`:
- Around line 29-46: Update the manual reseeding instructions in the SQL seed
comments so the listed DELETE statements are explicitly restricted to disposable
development databases and are not presented as a safe procedure for ordinary
local databases. Do not add arbitrary WHERE conditions; state that applying seed
changes to a database containing local data requires a separate seed-row
identification design first.
In `@src/test/java/com/DOCKin/ai/service/TranslateParallelBenchmarkTest.java`:
- Around line 46-55: Store the ExecutorService created for server.setExecutor in
a field, then update tearDown to shut it down and await its termination after
stopping the server. Preserve the existing server cleanup while ensuring the
fixed-thread-pool threads are fully terminated before the test completes.
In `@src/test/java/com/DOCKin/global/config/LocalMigrationDriftTest.java`:
- Around line 88-94: Update the migration validation flow in
LocalMigrationDriftTest to call flyway.validate() and explicitly collect entries
with MigrationState.FAILED from flyway.info().all(). Ensure checksum mismatches
and failed migrations cause the test to fail, while preserving the existing
version-filtered pending-migration handling.
In
`@src/test/java/com/DOCKin/global/validation/BeanValidationConstraintTest.java`:
- Around line 191-196: Update the class discovery logic around
BeanValidationConstraintTest to use Class.forName with initialization disabled,
collect each class name and loading failure cause instead of swallowing
Throwable, and fail the test when any discovered class cannot be loaded.
Preserve validation of successfully loaded classes while preventing incomplete
coverage from passing.
In `@src/test/java/com/DOCKin/global/web/PageableSortDefaultTest.java`:
- Around line 112-119: Update hasDefaultSort(Parameter) to inspect repeated
`@SortDefault` declarations via parameter.getAnnotationsByType(SortDefault.class)
rather than getAnnotation(SortDefault.class), and return true when any retrieved
annotation has a non-empty sort() value; preserve the existing PageableDefault
check.
In `@src/test/java/com/DOCKin/rag/chunking/SentenceLookbackMeasurementTest.java`:
- Around line 80-97: In SentenceLookbackMeasurementTest, add explicit
assertEquals checks that workLogAt120.forcedRatio() is zero before comparing
workLogAt400 results, and that longAt120.forcedRatio() is greater than zero
before asserting the reduction at longAt200. Add the required assertEquals
static import and keep the existing behavioral assertions unchanged.
In `@src/test/java/com/DOCKin/rag/service/AnnRecallMeasurementTest.java`:
- Around line 166-168: AnnRecallMeasurementTest의 세 측정 흐름에서 getConnection 호출만 별도의
SQLException 처리로 분리하고, 연결 실패일 때만 Assumptions.abort를 호출하세요. try-with-resources
전체를 감싸는 catch는 제거하거나 측정 단계와 분리하여 search, explainMs, measureWithPermissionFilter,
printPlan에서 발생한 SQLException이 테스트 실패로 전파되도록 수정하세요.
- Around line 519-553: Make session-setting restoration exception-safe in search
and explainMs by moving the reset of enable_indexscan into a finally block that
runs after the setting is changed, while preserving the existing query and
return behavior. Apply the same finally-based restoration pattern to the
hnsw.iterative_scan setting in the setup flow around the visible configuration
at lines 464-470, ensuring every modified connection setting is restored even
when SQL execution throws.
In `@src/test/java/com/DOCKin/rag/service/BruteForceSearchBenchmarkTest.java`:
- Around line 54-81: Separate connection establishment from the benchmark
execution in the test method containing the try-with-resources. Apply
Assumptions.abort only around DriverManager.getConnection failure; let
SQLException from prepareTable, measure, or dropTable propagate or otherwise
fail the test so SQL execution errors are not treated as skipped benchmarks.
In `@src/test/java/com/DOCKin/rag/service/CorpusSeedGeneratorTest.java`:
- Around line 40-44: Update CorpusSeedGeneratorTest so owner selectivity is an
explicit measurement variable rather than only implied by OWNER_COUNT, and
capture an EXPLAIN plan alongside every recall measurement. Ensure results using
idx_chunk_visibility or any non-HNSW plan are marked and excluded from HNSW
post-filter recall conclusions; keep the generated ownership distribution
aligned with the reported selectivity.
In `@src/test/java/com/DOCKin/rag/service/CrossLingualRetrievalTest.java`:
- Around line 94-98: Update CrossLingualRetrievalTest assertions to explicitly
validate recall@1 using vi.hits and en.hits, ensuring each result set contains
the expected number of correct top-ranked hits. If all five samples are
required, assert their hit counts equal KOREAN.length while retaining the
existing average-score checks.
---
Outside diff comments:
In `@src/main/java/com/DOCKin/ai/service/FastApiService.java`:
- Around line 124-153: saveTranslateLog에서 `@Transactional` 범위를 제거해 WorkLog 조회와
FastAPI 호출(Mono.zip(...).block())이 트랜잭션 밖에서 수행되도록 분리하십시오. 번역 응답을 받은 뒤 번역 행
upsert만 별도의 짧은 쓰기 트랜잭션 메서드로 위임하고, 동시 재번역 시 원자성이 유지되도록 기존 upsert 처리와 트랜잭션 경계를
보존하십시오.
In `@src/main/java/com/DOCKin/worklog/service/WorkLogsService.java`:
- Around line 40-69: Refactor createWorklog and createSttWorklog so validation
and WorkLog persistence occur in a short transaction that ends before invoking
uploadImage or processStt(...).block(); perform STT processing and S3 uploads
outside the transactional methods, then persist the resulting data in a separate
transaction, and delete any S3 objects created during the operation if saving
fails.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro Plus
Run ID: 9d324db2-2021-470b-860e-f456bc4533dc
⛔ Files ignored due to path filters (11)
app.jaris excluded by!**/*.jar,!**/*.jarmeasure-aws/app-night1.logis excluded by!**/*.logmeasure-aws/e1-20260812T164256/e1-sweep.csvis excluded by!**/*.csvmeasure-aws/e1-20260812T164256/run.logis excluded by!**/*.logmeasure-aws/e1-20260812T171704/e1-sweep.csvis excluded by!**/*.csvmeasure-aws/e1-20260812T171704/run.logis excluded by!**/*.logmeasure-aws/e8-20260812T164954/e8-samples.csvis excluded by!**/*.csvmeasure-aws/night1-20260812T164952/e1-sweep.csvis excluded by!**/*.csvmeasure-aws/night1-20260812T164952/sampler.logis excluded by!**/*.logmeasure-aws/night1-20260812T164952/sweep.logis excluded by!**/*.logmeasure-aws/night1.logis excluded by!**/*.log
📒 Files selected for processing (225)
.coderabbit.yaml.dockerignore.env.example.gitattributes.github/workflows/ci.yml.github/workflows/workflow.yml.gitignoreDockerfileREADME.mdbuild.gradlecompose.gc.yamlcompose.nocpu.yamlcompose.yamldocs/2026-07-04-checklist-domain-work-summary.mddocs/AWS-MEASUREMENT-PLAN.mddocs/AWS-MEASUREMENT-RESULTS.mddocs/JVM-GC-COLLECTORS.mddocs/PORTFOLIO-ROADMAP.mddocs/PROJECT-SCOPE.mddocs/SERVICE-SCALE-ASSUMPTIONS.mddocs/WORK-BACKLOG.mddocs/adr/0001-attendance-clockin-concurrency.mddocs/adr/0002-performance-improvement-backlog.mddocs/adr/0003-search-domain-expansion.mddocs/adr/0004-high-traffic-scaling-roadmap.mddocs/adr/0005-attendance-portfolio-gap-analysis.mddocs/adr/0006-rag-vector-search.mddocs/adr/0007-corpus-retention.mddocs/db/postgresql-schema.sqldocs/db/rebuild-hnsw-index.sqlgradlewinit.sqlmeasure-aws/e1-20260812T164256/env.mdmeasure-aws/e1-20260812T164256/raw/seq1-DOCKin-DB.txtmeasure-aws/e1-20260812T164256/raw/seq1-dockin-app.txtmeasure-aws/e1-20260812T164256/raw/seq1-dockin-embedding.txtmeasure-aws/e1-20260812T164256/raw/seq1-dockin-nginx.txtmeasure-aws/e1-20260812T164256/raw/seq1-dockin-redis.txtmeasure-aws/e1-20260812T164256/raw/seq2-DOCKin-DB.txtmeasure-aws/e1-20260812T164256/raw/seq2-dockin-app.txtmeasure-aws/e1-20260812T164256/raw/seq2-dockin-embedding.txtmeasure-aws/e1-20260812T164256/raw/seq2-dockin-nginx.txtmeasure-aws/e1-20260812T164256/raw/seq2-dockin-redis.txtmeasure-aws/e1-20260812T171704/env.mdmeasure-aws/e1-20260812T171704/raw/seq1-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq1-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq1-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq1-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq1-dockin-redis.txtmeasure-aws/e1-20260812T171704/raw/seq2-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq2-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq2-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq2-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq2-dockin-redis.txtmeasure-aws/e1-20260812T171704/raw/seq3-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq3-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq3-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq3-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq3-dockin-redis.txtmeasure-aws/e1-20260812T171704/raw/seq4-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq4-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq4-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq4-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq4-dockin-redis.txtmeasure-aws/e1-20260812T171704/raw/seq5-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq5-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq5-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq5-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq5-dockin-redis.txtmeasure-aws/e1-20260812T171704/raw/seq6-DOCKin-DB.txtmeasure-aws/e1-20260812T171704/raw/seq6-dockin-app.txtmeasure-aws/e1-20260812T171704/raw/seq6-dockin-embedding.txtmeasure-aws/e1-20260812T171704/raw/seq6-dockin-nginx.txtmeasure-aws/e1-20260812T171704/raw/seq6-dockin-redis.txtmeasure-aws/night1-20260812T164952/STATUSmeasure-aws/night1-20260812T164952/winner-cpusnginx/conf.d/default.confscripts/aws-bootstrap.shscripts/e1-tei-cpu-sweep.shscripts/e8-index-sampler.shscripts/night1-run.shsrc/main/java/com/DOCKin/DocKinSpringApplication.javasrc/main/java/com/DOCKin/absence/controller/AbsenceAdminController.javasrc/main/java/com/DOCKin/absence/controller/AbsenceRequestController.javasrc/main/java/com/DOCKin/absence/dto/AbsenceDecisionRequestDto.javasrc/main/java/com/DOCKin/absence/dto/AbsenceRequestCreateRequestDto.javasrc/main/java/com/DOCKin/absence/dto/AbsenceRequestResponseDto.javasrc/main/java/com/DOCKin/absence/event/AbsenceApprovedEvent.javasrc/main/java/com/DOCKin/absence/model/AbsenceRequest.javasrc/main/java/com/DOCKin/absence/model/AbsenceStatus.javasrc/main/java/com/DOCKin/absence/model/AbsenceType.javasrc/main/java/com/DOCKin/absence/repository/AbsenceRequestRepository.javasrc/main/java/com/DOCKin/absence/service/AbsenceRequestService.javasrc/main/java/com/DOCKin/ai/controller/AiController.javasrc/main/java/com/DOCKin/ai/dto/OnlineTranslateDomain.javasrc/main/java/com/DOCKin/ai/model/ChatHistory.javasrc/main/java/com/DOCKin/ai/model/ChatLog.javasrc/main/java/com/DOCKin/ai/model/TranslateLog.javasrc/main/java/com/DOCKin/ai/repository/TranslateRepository.javasrc/main/java/com/DOCKin/ai/service/FastApiService.javasrc/main/java/com/DOCKin/attendance/model/Attendance.javasrc/main/java/com/DOCKin/attendance/model/DayType.javasrc/main/java/com/DOCKin/attendance/model/WorkCalendar.javasrc/main/java/com/DOCKin/attendance/repository/AttendanceRepository.javasrc/main/java/com/DOCKin/attendance/repository/WorkCalendarRepository.javasrc/main/java/com/DOCKin/attendance/service/AbsenceApprovedListener.javasrc/main/java/com/DOCKin/attendance/service/AttendanceBatchService.javasrc/main/java/com/DOCKin/attendance/service/AttendanceService.javasrc/main/java/com/DOCKin/attendance/service/WorkCalendarService.javasrc/main/java/com/DOCKin/checklist/controller/ChecklistAdminController.javasrc/main/java/com/DOCKin/checklist/controller/ChecklistUserController.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistCheckRequestDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistCreateRequestDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistDetailResponseDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistItemRequestDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistItemSimpleResponseDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistItemStatusResponseDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistItemUpdateRequestDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistResponseDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistResultResponseDto.javasrc/main/java/com/DOCKin/checklist/dto/ChecklistUpdateRequestDto.javasrc/main/java/com/DOCKin/checklist/model/Checklist.javasrc/main/java/com/DOCKin/checklist/model/ChecklistItem.javasrc/main/java/com/DOCKin/checklist/model/ChecklistPhase.javasrc/main/java/com/DOCKin/checklist/model/ChecklistResult.javasrc/main/java/com/DOCKin/checklist/repository/ChecklistItemRepository.javasrc/main/java/com/DOCKin/checklist/repository/ChecklistRepository.javasrc/main/java/com/DOCKin/checklist/repository/ChecklistResultRepository.javasrc/main/java/com/DOCKin/checklist/service/ChecklistService.javasrc/main/java/com/DOCKin/checklist/service/ChecklistStatusService.javasrc/main/java/com/DOCKin/global/config/ClockConfig.javasrc/main/java/com/DOCKin/global/config/RedissonConfig.javasrc/main/java/com/DOCKin/global/config/WebClientConfig.javasrc/main/java/com/DOCKin/global/config/WebConfig.javasrc/main/java/com/DOCKin/global/config/WebSocketConfig.javasrc/main/java/com/DOCKin/global/error/ErrorCode.javasrc/main/java/com/DOCKin/global/error/GlobalExceptionHandler.javasrc/main/java/com/DOCKin/global/file/LogImage.javasrc/main/java/com/DOCKin/global/logging/TraceId.javasrc/main/java/com/DOCKin/global/logging/TraceIdFilter.javasrc/main/java/com/DOCKin/global/security/config/JacksonConfig.javasrc/main/java/com/DOCKin/global/security/config/SecurityConfig.javasrc/main/java/com/DOCKin/global/security/config/SecurityPathConfig.javasrc/main/java/com/DOCKin/global/security/jwt/JwtAuthFilter.javasrc/main/java/com/DOCKin/member/dto/MemberRequestDto.javasrc/main/java/com/DOCKin/member/model/Member.javasrc/main/java/com/DOCKin/member/model/WorkShift.javasrc/main/java/com/DOCKin/member/repository/MemberRepository.javasrc/main/java/com/DOCKin/member/service/MemberService.javasrc/main/java/com/DOCKin/rag/chunking/ChunkingStrategy.javasrc/main/java/com/DOCKin/rag/chunking/FixedSizeChunkingStrategy.javasrc/main/java/com/DOCKin/rag/dto/RetrievalResult.javasrc/main/java/com/DOCKin/rag/dto/RetrievedChunk.javasrc/main/java/com/DOCKin/rag/model/DocumentChunk.javasrc/main/java/com/DOCKin/rag/model/SourceType.javasrc/main/java/com/DOCKin/rag/model/Visibility.javasrc/main/java/com/DOCKin/rag/repository/ChunkVector.javasrc/main/java/com/DOCKin/rag/repository/DocumentChunkRepository.javasrc/main/java/com/DOCKin/rag/repository/NearestChunk.javasrc/main/java/com/DOCKin/rag/service/ChunkIndexWriter.javasrc/main/java/com/DOCKin/rag/service/EmbeddingClient.javasrc/main/java/com/DOCKin/rag/service/IndexingService.javasrc/main/java/com/DOCKin/rag/service/RagChatService.javasrc/main/java/com/DOCKin/rag/service/RetrievalService.javasrc/main/java/com/DOCKin/rag/service/SeedIndexingRunner.javasrc/main/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusRequestDto.javasrc/main/java/com/DOCKin/safetyCourse/model/SafetyCourse.javasrc/main/java/com/DOCKin/safetyCourse/repository/SafetyCourseRepository.javasrc/main/java/com/DOCKin/worklog/controller/WorkLogsController.javasrc/main/java/com/DOCKin/worklog/dto/WorkLogDto.javasrc/main/java/com/DOCKin/worklog/dto/Work_logsDto.javasrc/main/java/com/DOCKin/worklog/model/Comment.javasrc/main/java/com/DOCKin/worklog/model/WorkLog.javasrc/main/java/com/DOCKin/worklog/model/WorkLogImage.javasrc/main/java/com/DOCKin/worklog/model/Work_logs.javasrc/main/java/com/DOCKin/worklog/repository/WorkLogRepository.javasrc/main/java/com/DOCKin/worklog/repository/Work_logsRepository.javasrc/main/java/com/DOCKin/worklog/service/CommentService.javasrc/main/java/com/DOCKin/worklog/service/WorkLogsService.javasrc/main/resources/application-seed.propertiessrc/main/resources/application.propertiessrc/main/resources/data.sqlsrc/main/resources/db/migration/V1__pgvector_and_document_chunks.sqlsrc/main/resources/db/migration/V2__baseline_existing_tables.sqlsrc/main/resources/db/migration/V3__work_log_child_fk_indexes.sqlsrc/main/resources/db/migration/V4__users_child_fk_indexes.sqlsrc/main/resources/db/migration/V5__work_logs_user_id_not_null.sqlsrc/main/resources/db/seed/R__seed_sample.sqlsrc/main/resources/schema.sqlsrc/main/resources/templates/chat_test.htmlsrc/test/java/com/DOCKin/DOCKin_spring/DocKinSpringApplicationTests.javasrc/test/java/com/DOCKin/absence/service/AbsenceRequestServiceTest.javasrc/test/java/com/DOCKin/absence/service/LeaveBalanceConcurrencyTest.javasrc/test/java/com/DOCKin/ai/service/TranslateParallelBenchmarkTest.javasrc/test/java/com/DOCKin/attendance/service/AbsenceApprovedListenerTest.javasrc/test/java/com/DOCKin/attendance/service/AttendanceBatchServiceTest.javasrc/test/java/com/DOCKin/attendance/service/AttendanceServiceTest.javasrc/test/java/com/DOCKin/attendance/service/WorkCalendarServiceTest.javasrc/test/java/com/DOCKin/checklist/service/ChecklistServiceTest.javasrc/test/java/com/DOCKin/checklist/service/ChecklistStatusServiceTest.javasrc/test/java/com/DOCKin/global/config/ActuatorEndpointTest.javasrc/test/java/com/DOCKin/global/config/LocalMigrationDriftTest.javasrc/test/java/com/DOCKin/global/config/SchemaValidationTest.javasrc/test/java/com/DOCKin/global/config/SeedDataTest.javasrc/test/java/com/DOCKin/global/error/GlobalExceptionHandlerTest.javasrc/test/java/com/DOCKin/global/logging/TraceIdFilterTest.javasrc/test/java/com/DOCKin/global/testsupport/ContainerTestSupport.javasrc/test/java/com/DOCKin/global/validation/BeanValidationConstraintTest.javasrc/test/java/com/DOCKin/global/web/PageableSortDefaultTest.javasrc/test/java/com/DOCKin/member/repository/LockTimeoutVerificationTest.javasrc/test/java/com/DOCKin/rag/chunking/FixedSizeChunkingStrategyTest.javasrc/test/java/com/DOCKin/rag/chunking/SentenceLookbackMeasurementTest.javasrc/test/java/com/DOCKin/rag/repository/FlywayMigrationTest.javasrc/test/java/com/DOCKin/rag/repository/HibernateBatchInsertVerificationTest.javasrc/test/java/com/DOCKin/rag/repository/VectorSearchQueryTest.javasrc/test/java/com/DOCKin/rag/repository/VectorTypeMappingTest.javasrc/test/java/com/DOCKin/rag/service/AnnRecallMeasurementTest.javasrc/test/java/com/DOCKin/rag/service/BruteForceSearchBenchmarkTest.javasrc/test/java/com/DOCKin/rag/service/CorpusSeedGeneratorTest.javasrc/test/java/com/DOCKin/rag/service/CrossLingualRetrievalTest.javasrc/test/java/com/DOCKin/rag/service/EmbeddingClientTest.javasrc/test/java/com/DOCKin/rag/service/RetrievalServiceTest.javasrc/test/java/com/DOCKin/safetyCourse/dto/SafetyWatchStatusValidationTest.javasrc/test/java/com/DOCKin/worklog/service/WorkLogListBenchmarkTest.javasrc/test/java/com/DOCKin/worklog/service/WorkLogListQueryCountTest.java
💤 Files with no reviewable changes (9)
- .github/workflows/workflow.yml
- src/main/resources/data.sql
- src/main/java/com/DOCKin/global/file/LogImage.java
- src/main/java/com/DOCKin/worklog/model/Work_logs.java
- init.sql
- src/main/java/com/DOCKin/worklog/dto/Work_logsDto.java
- src/main/java/com/DOCKin/ai/model/ChatHistory.java
- src/main/java/com/DOCKin/DocKinSpringApplication.java
- src/main/java/com/DOCKin/worklog/repository/Work_logsRepository.java
| //s3 upload | ||
| implementation 'org.springframework.cloud:spring-cloud-starter-aws:2.2.6.RELEASE' |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | 🏗️ Heavy lift
🧩 Analysis chain
🌐 Web query:
Is org.springframework.cloud:spring-cloud-starter-aws:2.2.6.RELEASE compatible with Spring Boot 4.0.1 and Spring Framework 7?
💡 Result:
No, org.springframework.cloud:spring-cloud-starter-aws:2.2.6.RELEASE is not compatible with Spring Boot 4.0.1 or Spring Framework 7 [1][2][3]. Spring Cloud AWS 2.2.6.RELEASE is a legacy version originally released in 2021 [4] and is designed for much older versions of the Spring ecosystem (specifically those aligned with Spring Cloud Hoxton/Spring Boot 2.2.x–2.3.x) [5][6]. It is neither supported nor compatible with modern Spring Boot 3.x or 4.x releases [7]. To use Spring Cloud AWS with Spring Boot 4.0.x and Spring Framework 7.x, you must upgrade to the latest major version of the library (Spring Cloud AWS 4.x) [1][2][3]. Versions of Spring Cloud AWS 4.x and higher are specifically maintained to provide full compatibility with Spring Boot 4.0.x and Spring Framework 7.x [1][8].
Citations:
- 1: https://github.com/awspring/spring-cloud-aws
- 2: https://awspring.io/what-is-spring-cloud-aws
- 3: https://github.com/awspring/spring-cloud-aws/blob/main/README.md
- 4: https://repo1.maven.org/maven2/org/springframework/cloud/spring-cloud-starter-aws/2.2.6.RELEASE/
- 5: https://docs.spring.io/spring-cloud-aws/docs/2.2.6.RELEASE/2.2.6.RELEASE/reference/html/
- 6: https://spring.io/blog/2021/03/17/spring-cloud-aws-2-3-is-now-available
- 7: https://stackoverflow.com/questions/76172912/micrometer-stopped-publishing-metrics-after-migration-to-spring-boot-3
- 8: https://github.com/awspring/spring-cloud-aws/releases
🏁 Script executed:
#!/bin/bash
set -e
printf '%s\n' '--- build.gradle dependency context ---'
sed -n '1,150p' build.gradle
printf '%s\n' '--- Spring Boot and Spring Cloud AWS references ---'
rg -n --glob '!build/**' --glob '!node_modules/**' \
'spring-boot|spring-cloud|spring-cloud-starter-aws|awspring|spring-framework' .
printf '%s\n' '--- Maven metadata and POM for the declared artifact ---'
python3 - <<'PY'
import urllib.request
import xml.etree.ElementTree as ET
base = "https://repo1.maven.org/maven2/org/springframework/cloud/spring-cloud-starter-aws/2.2.6.RELEASE"
for name in ("spring-cloud-starter-aws-2.2.6.RELEASE.pom",):
url = f"{base}/{name}"
print(f"URL: {url}")
data = urllib.request.urlopen(url, timeout=20).read()
root = ET.fromstring(data)
ns = {"m": "http://maven.apache.org/POM/4.0.0"}
for tag in ("parent", "properties", "dependencies"):
node = root.find(f"m:{tag}", ns)
if node is not None:
print(ET.tostring(node, encoding="unicode")[:5000])
PY
PY_STATUS=$?
exit "$PY_STATUS"Repository: DOCKin-project/DOCKin-backend
Length of output: 38499
🌐 Web query:
official Spring Cloud AWS compatibility matrix Spring Boot 4 Spring Framework 7 spring-cloud-starter-aws 2.2.6.RELEASE
💡 Result:
According to the official Spring Cloud AWS compatibility documentation, Spring Cloud AWS version 4.x.x is the version compatible with Spring Boot 4.0.x and Spring Framework 7.0.x [1][2][3]. Regarding spring-cloud-starter-aws 2.2.6.RELEASE, it is an older version and is not compatible with Spring Boot 4 or Spring Framework 7. The 2.2.x series of Spring Cloud AWS was designed for significantly older versions of the Spring ecosystem (specifically Spring Boot 2.x and Spring Framework 5.x) [1][3]. For projects upgrading to Spring Boot 4 and Spring Framework 7, you must migrate to Spring Cloud AWS 4.x.x [4]. Additionally, please note that the major versions of Spring Cloud AWS (e.g., 2.x vs 3.x/4.x) contain significant architectural changes, including the migration to the AWS SDK for Java 2.x and shifts in supported AWS services [3].
Citations:
- 1: https://github.com/awspring/spring-cloud-aws?tab=readme-ov-file
- 2: https://github.com/awspring/spring-cloud-aws/blob/main/README.md
- 3: https://awspring.io/what-is-spring-cloud-aws
- 4: https://github.com/awspring/spring-cloud-aws/releases
🌐 Web query:
site:docs.awspring.io/spring-cloud-aws Spring Cloud AWS 4 Spring Boot 4 compatibility
💡 Result:
As of August 13, 2026, there is no official release of Spring Boot 4 [1][2][3]. Because Spring Boot 4 does not exist, there is no formal compatibility status between Spring Cloud AWS 4 and Spring Boot 4. Spring Cloud AWS version 4.x is designed to integrate with established versions of Spring Boot (typically Spring Boot 3.x) [1][2]. You can verify the specific Spring Boot dependencies supported by a particular version of Spring Cloud AWS by checking the Bill of Materials (BOM) for that release in the official Spring Cloud AWS documentation [3]. To confirm the versions supported in your project, consult the Spring Cloud AWS reference documentation for the specific version you are using [1][3].
Citations:
- 1: https://docs.awspring.io/spring-cloud-aws/docs/4.0.2/reference/html/index.html
- 2: https://docs.awspring.io/spring-cloud-aws/docs/4.0.0/reference/html/index.html
- 3: https://docs.awspring.io/spring-cloud-aws/docs/3.4.0/reference/html/index.html
🌐 Web query:
site:github.com/awspring/spring-cloud-aws spring-cloud-aws-starter-s3 Spring Boot 4
💡 Result:
Spring Cloud AWS 4.x is the version line designed for compatibility with Spring Boot 4.0 and Spring Framework 7.0 [1][2]. The starter for S3 is available as spring-cloud-aws-starter-s3, which is included in the Spring Cloud AWS 4.x releases [1][2]. As of today, August 13, 2026, the current stable releases in this series include version 4.0.1 [3]. You can use this starter by adding the corresponding dependency for Spring Cloud AWS 4.x to your project [1][3]. Key details regarding this compatibility: - Version alignment: Spring Cloud AWS 4.0.0 and subsequent 4.x versions are explicitly aligned with the Spring Boot 4.x ecosystem [1][2]. - S3 Integration: The S3 starter (spring-cloud-aws-starter-s3) is part of this release and includes integrated support for Spring Integration AWS [1][2]. - Official Documentation: You can find detailed reference documentation for the current version at the official Spring Cloud AWS documentation site [3]. It is recommended to use the latest stable 4.x release (4.0.1 or later) to ensure full compatibility with the Spring Boot 4.x release [2][3].
Citations:
- 1: https://github.com/awspring/spring-cloud-aws/releases/tag/v4.0.0
- 2: https://github.com/awspring/spring-cloud-aws/releases
- 3: https://github.com/awspring/spring-cloud-aws/releases/tag/v4.0.1
🏁 Script executed:
#!/bin/bash
set -e
fetch() {
url="$1"
echo "--- $url ---"
curl -k -fsSL --max-time 20 "$url"
echo
}
fetch 'https://repo1.maven.org/maven2/org/springframework/cloud/spring-cloud-starter-aws/2.2.6.RELEASE/spring-cloud-starter-aws-2.2.6.RELEASE.pom' \
| python3 -c '
import sys, xml.etree.ElementTree as ET
root=ET.parse(sys.stdin).getroot()
ns={"m":"http://maven.apache.org/POM/4.0.0"}
for section in ("parent","properties","dependencies"):
node=root.find("m:"+section, ns)
if node is not None:
print(ET.tostring(node, encoding="unicode"))
'
fetch 'https://api.github.com/repos/awspring/spring-cloud-aws/releases?per_page=10' \
| python3 -c '
import json,sys
for r in json.load(sys.stdin):
print(r.get("tag_name"), r.get("published_at"), r.get("name"))
'
fetch 'https://raw.githubusercontent.com/awspring/spring-cloud-aws/main/README.md' \
| rg -n -i -C 3 'spring boot|spring framework|compatib|4\.0|starter-s3' || trueRepository: DOCKin-project/DOCKin-backend
Length of output: 533
🌐 Web query:
site:github.com/awspring/spring-cloud-aws/releases/tag "Spring Boot 4" "Spring Cloud AWS"
💡 Result:
As of August 13, 2026, Spring Cloud AWS version 4.x is the current major release line designed specifically for compatibility with Spring Boot 4.x and Spring Framework 7.x [1][2]. Key details regarding Spring Cloud AWS 4.x and Spring Boot 4: Compatibility: Spring Cloud AWS 4.0.0 and subsequent releases (such as 4.0.1 and 4.0.2) provide full support for the Spring Boot 4 ecosystem [1][3][4]. Release Timeline: The stable version 4.0.0 was released on January 28, 2026 [1]. Major Changes: This version introduces significant architectural updates, most notably the integration of the Spring Integration AWS project directly into the Spring Cloud AWS codebase, unifying AWS-related integration support within the ecosystem [1][2]. Security and Maintenance: Recent updates, such as version 4.0.2 (released April 30, 2026), have introduced enhancements like mandatory SNS message signature verification by default [4]. Developers using Spring Boot 4 should utilize the 4.x line of Spring Cloud AWS to ensure full compatibility with the Spring Framework 7 baseline [1][3].
Citations:
- 1: https://github.com/awspring/spring-cloud-aws/releases/tag/v4.0.0
- 2: https://github.com/awspring/spring-cloud-aws/releases/tag/v4.0.0-M1
- 3: https://github.com/awspring/spring-cloud-aws/releases/tag/v4.0.1
- 4: https://github.com/awspring/spring-cloud-aws/releases/tag/v4.0.2
Spring Boot 4 호환 Spring Cloud AWS 버전으로 마이그레이션하십시오.
org.springframework.cloud:spring-cloud-starter-aws:2.2.6.RELEASE는 Spring Boot 2.x용 레거시 의존성입니다. org.springframework.boot:4.0.1과 호환되지 않아 자동 구성 또는 S3 클라이언트 초기화가 실패할 수 있습니다. 공식 호환성 매트릭스에 맞는 starter와 버전으로 교체하고 API 변경을 반영하십시오.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@build.gradle` around lines 106 - 107, Update the S3 dependency in the build
configuration from the legacy spring-cloud-starter-aws 2.2.6.RELEASE to the
Spring Cloud AWS starter and version compatible with Spring Boot 4.0.1,
following the official compatibility matrix. Then adapt the S3 client
initialization and related APIs to the selected starter’s current API.
| - RBAC은 프로젝트 관례대로(`@PreAuthorize` 미사용) 서비스 레이어 수동 체크(`SafetyCourseService` 패턴). | ||
| - 관리자(`ChecklistAdminController`, `/api/checklist/admin`)와 사용자(`ChecklistUserController`, `/api/checklist/user`) 컨트롤러 분리. | ||
| - `ErrorCode`에 `CK001~CK007` 추가, `schema.sql` 갱신(컬럼명 `role`→`phase`, `checklist_results`를 `checklist_item_id` 참조로 변경, 중복 `equipment_id` 컬럼 제거). | ||
| - 단위 테스트: `ChecklistServiceTest`(12개), `ChecklistStatusServiceTest`(8개) — 전부 통과. |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
수치의 실측 여부와 근거를 명시하십시오. 현재 수치는 테스트 결과 또는 용량 추정이지만, 문서 규약의 측정 상태 표기가 일관되지 않습니다.
docs/2026-07-04-checklist-domain-work-summary.md#L17-L17: 테스트 실행 시점, 실행 명령 또는 CI 산출물 경로를 추가하십시오.docs/AWS-MEASUREMENT-PLAN.md#L81-L81: 780MB 및 390MB 추정에 산정 근거와[실측 필요]를 추가하십시오.docs/AWS-MEASUREMENT-RESULTS.md#L90-L93: 40MB, 50MB, 300MB 추정에 원본 측정 경로를 연결하거나[실측 필요]를 추가하십시오.
As per path instructions, "가정에 [실측 필요], 미검증 설정에 [미검증]을 붙이는 규칙"을 적용해야 합니다.
📍 Affects 3 files
docs/2026-07-04-checklist-domain-work-summary.md#L17-L17(this comment)docs/AWS-MEASUREMENT-PLAN.md#L81-L81docs/AWS-MEASUREMENT-RESULTS.md#L90-L93
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@docs/2026-07-04-checklist-domain-work-summary.md` at line 17, 문서의 측정 상태 표기를
근거와 함께 일관되게 보완하십시오. docs/2026-07-04-checklist-domain-work-summary.md 17-17의 테스트
수치에 실행 시점과 실행 명령 또는 CI 산출물 경로를 추가하고, docs/AWS-MEASUREMENT-PLAN.md 81-81의
780MB·390MB 추정에는 산정 근거와 [실측 필요]를 표시하십시오. docs/AWS-MEASUREMENT-RESULTS.md 90-93의
40MB·50MB·300MB 추정에는 원본 측정 경로를 연결하거나 [실측 필요]를 추가하며, 가정에는 [실측 필요], 미검증 설정에는 [미검증]
표기를 적용하십시오.
Source: Path instructions
| } catch (SQLException e) { | ||
| Assumptions.abort("PostgreSQL 접속 실패로 건너뜁니다: " + e.getMessage()); | ||
| } |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
측정 중 발생한 SQLException이 "접속 실패"로 삼켜져 테스트가 건너뜀으로 보고된다.
라인 166의 catch는 라인 119의 try-with-resources 블록 전체를 감싼다. getConnection의 실패만 잡는 것이 아니다. search, explainMs, measureWithPermissionFilter, printPlan 중 어디서 나는 SQLException도 여기로 온다. 그러면 Assumptions.abort가 호출되어 결과가 건너뜀이 되고, 출력 메시지는 사실과 다른 "PostgreSQL 접속 실패"가 된다.
라인 225와 라인 290도 같다. 접속 단계만 별도로 판정한다.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/test/java/com/DOCKin/rag/service/AnnRecallMeasurementTest.java` around
lines 166 - 168, AnnRecallMeasurementTest의 세 측정 흐름에서 getConnection 호출만 별도의
SQLException 처리로 분리하고, 연결 실패일 때만 Assumptions.abort를 호출하세요. try-with-resources
전체를 감싸는 catch는 제거하거나 측정 단계와 분리하여 search, explainMs, measureWithPermissionFilter,
printPlan에서 발생한 SQLException이 테스트 실패로 전파되도록 수정하세요.
| private Set<Long> search(Connection conn, String vec, String where, boolean useIndex) throws SQLException { | ||
| Set<Long> ids = new LinkedHashSet<>(); | ||
| try (Statement st = conn.createStatement()) { | ||
| st.execute("SET enable_indexscan = " + (useIndex ? "on" : "off")); | ||
| try (ResultSet rs = st.executeQuery(sql(vec, where))) { | ||
| while (rs.next()) { | ||
| ids.add(rs.getLong(1)); | ||
| } | ||
| } | ||
| st.execute("SET enable_indexscan = on"); | ||
| } | ||
| return ids; | ||
| } | ||
|
|
||
| private double explainMs(Connection conn, String vec, String where, boolean useIndex) throws SQLException { | ||
| try (Statement st = conn.createStatement()) { | ||
| st.execute("SET enable_indexscan = " + (useIndex ? "on" : "off")); | ||
| // 워밍 1회 후 측정 — 첫 실행은 캐시 미스가 섞인다. | ||
| st.executeQuery(sql(vec, where)).close(); | ||
| double best = Double.MAX_VALUE; | ||
| for (int i = 0; i < 3; i++) { | ||
| try (ResultSet rs = st.executeQuery("EXPLAIN (ANALYZE, TIMING ON) " + sql(vec, where))) { | ||
| while (rs.next()) { | ||
| String line = rs.getString(1); | ||
| if (line.startsWith("Execution Time:")) { | ||
| best = Math.min(best, Double.parseDouble( | ||
| line.replace("Execution Time:", "").replace("ms", "").trim())); | ||
| } | ||
| } | ||
| } | ||
| } | ||
| st.execute("SET enable_indexscan = on"); | ||
| return best == Double.MAX_VALUE ? -1 : best; | ||
| } | ||
| } |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
SET enable_indexscan 복원이 예외 안전하지 않다. 세션이 오염되면 이후 숫자가 조용히 틀린다.
라인 522는 enable_indexscan을 off로 바꾸고, 라인 528이 그것을 on으로 되돌린다. 라인 523의 executeQuery가 던지면 라인 528은 실행되지 않는다. try-with-resources는 Statement만 닫고 세션 설정은 되돌리지 않는다. explainMs의 라인 535와 라인 550도 같다.
같은 Connection을 계속 쓰므로 오염이 이어진다. 라인 318과 라인 471은 SQLException을 잡고 측정을 계속하므로 그 경로에서 실제로 이어진다. 그다음부터 useIndex=true로 부른 측정이 순차 스캔으로 실행되고, avgRecall은 1.000을 낸다. 라인 347-348의 주석이 경계한 바로 그 오독이 데이터 안에서 일어난다.
finally로 되돌린다. 라인 464-470의 hnsw.iterative_scan도 같은 처리가 필요하다.
🛡️ 제안 수정
private Set<Long> search(Connection conn, String vec, String where, boolean useIndex) throws SQLException {
Set<Long> ids = new LinkedHashSet<>();
try (Statement st = conn.createStatement()) {
st.execute("SET enable_indexscan = " + (useIndex ? "on" : "off"));
- try (ResultSet rs = st.executeQuery(sql(vec, where))) {
- while (rs.next()) {
- ids.add(rs.getLong(1));
- }
- }
- st.execute("SET enable_indexscan = on");
+ try {
+ try (ResultSet rs = st.executeQuery(sql(vec, where))) {
+ while (rs.next()) {
+ ids.add(rs.getLong(1));
+ }
+ }
+ } finally {
+ // 되돌리지 못하면 이후의 'ANN' 측정이 전부 순차 스캔이 된다.
+ st.execute("RESET enable_indexscan");
+ }
}
return ids;
}🧰 Tools
🪛 OpenGrep (1.26.0)
[ERROR] 522-522: SQL query built via string concatenation passed to Statement.execute*(). Use PreparedStatement with parameterized queries instead.
(coderabbit.sql-injection.java-statement-concat)
[ERROR] 535-535: SQL query built via string concatenation passed to Statement.execute*(). Use PreparedStatement with parameterized queries instead.
(coderabbit.sql-injection.java-statement-concat)
[ERROR] 540-540: SQL query built via string concatenation passed to Statement.execute*(). Use PreparedStatement with parameterized queries instead.
(coderabbit.sql-injection.java-statement-concat)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/test/java/com/DOCKin/rag/service/AnnRecallMeasurementTest.java` around
lines 519 - 553, Make session-setting restoration exception-safe in search and
explainMs by moving the reset of enable_indexscan into a finally block that runs
after the setting is changed, while preserving the existing query and return
behavior. Apply the same finally-based restoration pattern to the
hnsw.iterative_scan setting in the setup flow around the visible configuration
at lines 464-470, ensuring every modified connection setting is restored even
when SQL execution throws.
| try (Connection conn = DriverManager.getConnection(URL, "root", password)) { | ||
| System.out.println(); | ||
| System.out.println("=== 브루트포스 검색 실측 (PostgreSQL / " + DIM + "차원 / top-" + TOP_K + ") ==="); | ||
| System.out.printf("%-10s | %-22s | %-10s | %-11s | %-11s | %-11s | %-9s%n", | ||
| "적재 청크", "경로", "후보 수", "조회(ms)", "계산(ms)", "합계(ms)", "벡터(MB)"); | ||
| System.out.println("-".repeat(105)); | ||
|
|
||
| for (int scale : SCALES) { | ||
| prepareTable(conn, scale); | ||
|
|
||
| // 두 경로를 모두 잰다. | ||
| // - 일반 사용자: 권한 선필터가 후보를 크게 줄인다(현실적인 평균 경로) | ||
| // - 관리자: 필터가 없어 전체를 스캔한다(최악 케이스 = Phase 2 전환 판단 기준) | ||
| for (boolean admin : new boolean[]{false, true}) { | ||
| Result result = measure(conn, admin); | ||
| System.out.printf("%-10s | %-22s | %-10s | %-11.1f | %-11.1f | %-11.1f | %-9.1f%n", | ||
| String.format("%,d", scale), | ||
| admin ? "관리자(전체 스캔)" : "일반 사용자(선필터)", | ||
| String.format("%,d", result.candidates), | ||
| result.fetchMs, result.computeMs, result.totalMs, result.vectorMb); | ||
| } | ||
| } | ||
| System.out.println(); | ||
|
|
||
| dropTable(conn); | ||
| } catch (SQLException e) { | ||
| Assumptions.abort("PostgreSQL(localhost:5432)에 접속할 수 없어 벤치마크를 건너뜁니다: " + e.getMessage()); | ||
| } |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
SQL 실행 실패를 연결 실패로 처리하면 안 됩니다.
Line 79는 prepareTable, measure, dropTable의 모든 SQLException을 skip으로 바꿉니다. 테이블 정의 오류나 조회 오류도 측정 건너뜀으로 처리되어 실패한 벤치마크가 통과합니다.
DriverManager.getConnection 실패만 Assumptions.abort로 처리하십시오. 연결 후 발생한 SQL 오류는 테스트를 실패시켜야 합니다.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/test/java/com/DOCKin/rag/service/BruteForceSearchBenchmarkTest.java`
around lines 54 - 81, Separate connection establishment from the benchmark
execution in the test method containing the try-with-resources. Apply
Assumptions.abort only around DriverManager.getConnection failure; let
SQLException from prepareTable, measure, or dropTable propagate or otherwise
fail the test so SQL execution errors are not treated as skipped benchmarks.
| * <p><b>1. 소유자를 3명에서 {@value #OWNER_COUNT}명으로.</b> 이전에는 | ||
| * {@code R__seed_sample.sql}이 만든 {@code worker01~03}에 분배했고, 그래서 각자가 코퍼스의 | ||
| * 약 1/3을 가졌다. 권한 선필터의 선택도가 33.6%라는 뜻인데, <b>HNSW가 필터에서 무너지는 것은 | ||
| * 선택도가 높을수록 심해지므로 이것은 애초에 약한 조건이다</b>(ADR-0006 8-2). | ||
| * 실제 서비스는 사용자가 수백 명이고 각자 1% 미만을 소유한다. 여기서는 0.2%로 만든다. |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
선택도 조건과 실제 실행 계획을 함께 검증하세요.
OWNER_COUNT = 500은 소유자 선택도를 약 0.2%로 만듭니다. 그러나 현재 측정 기록에서 0.34% 조건은 HNSW가 아니라 idx_chunk_visibility를 사용했습니다. 이 생성기만으로는 HNSW 후필터의 recall 저하를 검증할 수 없습니다.
선택도를 측정 변수로 만들고, 각 recall 결과에 EXPLAIN 계획을 함께 저장하세요. HNSW를 사용하지 않은 결과는 HNSW recall 결과로 해석하지 마세요.
As per path instructions, “선언과 실제가 어긋나는 곳을 우선 본다” and “이 설정/애노테이션이 실제로 효과가 있는가를 물어라.”
Also applies to: 95-102
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/test/java/com/DOCKin/rag/service/CorpusSeedGeneratorTest.java` around
lines 40 - 44, Update CorpusSeedGeneratorTest so owner selectivity is an
explicit measurement variable rather than only implied by OWNER_COUNT, and
capture an EXPLAIN plan alongside every recall measurement. Ensure results using
idx_chunk_visibility or any non-HNSW plan are marked and excluded from HNSW
post-filter recall conclusions; keep the generated ownership distribution
aligned with the reported selectivity.
Source: Path instructions
| // 교차언어가 성립하려면 최소한 정답이 오답보다 일관되게 가까워야 한다. | ||
| assertTrue(vi.correctAvg > vi.wrongAvg, | ||
| "베트남어 질의에서 정답 원문이 오답보다 가깝지 않다 - 교차언어가 성립하지 않는다"); | ||
| assertTrue(en.correctAvg > en.wrongAvg, | ||
| "영어 질의에서 정답 원문이 오답보다 가깝지 않다"); |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
recall@1을 직접 검증해야 합니다.
현재 assertion은 평균 점수만 비교합니다. 모든 질의에서 오답이 1위여도 정답 평균 점수가 전체 오답 평균보다 높으면 테스트가 통과할 수 있습니다.
vi.hits와 en.hits에 명시적 기준을 적용하십시오. 이 5개 표본이 모두 필수라면 KOREAN.length와 같은지 검증하십시오.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/test/java/com/DOCKin/rag/service/CrossLingualRetrievalTest.java` around
lines 94 - 98, Update CrossLingualRetrievalTest assertions to explicitly
validate recall@1 using vi.hits and en.hits, ensuring each result set contains
the expected number of correct top-ranked hits. If all five samples are
required, assert their hit counts equal KOREAN.length while retaining the
existing average-score checks.
무엇을 고쳤나
2026-08-13 밤 1 무인 측정(E1/E8)이 코퍼스의 13%에서 끝났는데
STATUS에는OK가 남았다. 그 오판을 만든 자리 셋을 고치고, 밤 1의 결과를 문서로 남긴다.1. 앱이 완주와 중단을 다른 줄로 찍는다 (
e2e2e4c)indexAll()은 예외를 잡아 커밋된 분량을 살린 뒤 "인덱싱 종료"를 찍는다. 완주한 실행과 죽은 실행이 같은 문자열로 끝났다.로그 문자열만 가르면 거짓말이 절반 남는다.
int를 돌려주는 한 호출자는 완주를 판정할 수 없어서 — 중단돼도 그때까지의 건수가 그대로 나온다 —SeedIndexingRunner도 죽은 실행에 "색인 완료"를 찍고 있었다. 반환형을IndexRun(embedded, totalChunks, elapsedMs, abortReason)으로 바꿨다.2. 무인 실행의 두 판정 (
f4ee899)인덱싱 중단이 보이면 그 자리에서 실패로 끝내고 사유를STATUS에 옮긴다. 끝난 이유(end_reason)도 남긴다sort -n | head -1이 첫 줄을 집었다. 최고값의 +TIE_PCT%(기본 2) 안을 동률로 보고 그중 가장 낮은 상한을 고른다.1c0d98f가 이미 의도로 적어둔 규칙인데 구현이 하지 않고 있었다awk의n이 초기화 전에 첨자로 쓰여 첫 행이ms[""]에 들어가고ms[0]이 유령 0행이 되던 것을 고쳤다. 실전 CSV에서 답이 맞았던 건 우연이었다3. 밤 1 결과 기록 (
504e6af)docs/AWS-MEASUREMENT-RESULTS.md를 새로 만들었다. 계획은AWS-MEASUREMENT-PLAN.md, 결과는 밤별로 여기에 쌓는다. 산출물(measure-aws/, CSV·로그)도 함께 넣었다.밤 1이 실제로 얻은 것
E1 — 결론이 났다. TEI 상한 2.0~6.0에서 처리량이 전부 62.6ms/청크, 75.0ms/원본으로 동일하다. 계획 E1의 성공 판정 "상한을 올려도 원본당 ms가 안 변하면 병목이 TEI가 아니다" 그대로다. cgroup
cpu.max직접 확인, 통제군(마지막 2.0) 일치, 조건마다 같은 문서 재색인, 앱 로그의 조건별 투입량 6회 모두 4,907건 — 네 겹으로 받쳐진다.E8 — 미완. 0→25,815청크 구간은 75.0ms/원본으로 평평하다. 다만 26,285청크면
shared_buffers128MB 안에 전부 앉으므로, 보려던 열화는 멈춘 지점 이후에 시작된다.남은 것
색인 상호배제는 아직 없다. cron과 기동 러너가 같은 행을 집으면
uk_chunk_source로 터진다. 이 PR이 고치는 것은 "그것이 일어났을 때 알 수 있는가"까지다.검증
compileJava/compileTestJavaexit 0bash -n scripts/night1-run.sh통과app-night1.log에 걸어 사유 추출 확인🤖 Generated with Claude Code
Summary by CodeRabbit
새 기능
개선