Apache Arrow Dataset in Practice: The Complete Guide for Developers and Engineers
- 저자
- 출판사
- 언어학습
- 영어
- 형식
- 컬렉션
논픽션
"Apache Arrow Dataset in Practice"
"Apache Arrow Dataset in Practice" is a comprehensive guide for data engineers, analysts, and systems architects seeking to master high-performance, cross-language in-memory analytics using the Apache Arrow ecosystem. This authoritative book begins by setting the stage with a rich overview of Arrow’s evolution in the context of modern data interchange, deeply exploring its columnar in-memory format, abstractions like schemas and record batches, and the Dataset API's foundational principles. By blending theory with hands-on design philosophy and performance motivations, the introduction thoroughly prepares readers to leverage Arrow’s full potential in contemporary data workflows.
The heart of the book delves deeply into practical applications, covering sophisticated aspects of the Dataset API, including storage layer integration, partitioning, schema management, and expression-based filtering for scalable analytics. Readers learn efficient ingestion strategies, rigorous data validation techniques, vectorized transformations, and robust error handling to maintain data quality from source to export. Advanced chapters illuminate the mechanics of query processing—from vectorized execution and predicate pushdown to handling complex data types, aggregations, and performant joins—equipping practitioners with tools to optimize analytic workloads at any scale.
Beyond core functionalities, the book dedicates thorough coverage to real-world operations: achieving scalability across distributed environments, integrating seamlessly with leading analytics engines and data science toolkits, and maintaining security, privacy, and compliance throughout the data lifecycle. Practical guidance on debugging, optimization, and cost control is matched with a forward-looking perspective on extending Arrow and engaging with its vibrant open-source community. Through detailed case studies and in-depth technical advice, "Apache Arrow Dataset in Practice" stands as an indispensable resource for building next-generation, interoperable data applications.
© 2025 HiTeX Press (전자책): 6610000964550
출시일
전자책: 2025년 7월 12일
다른 사람들도 즐겼습니다 ...
- 어린이날이 사라진다고? 노수미
- 그 많던 싱아는 누가 다 먹었을까 박완서
- 신기한 맛 도깨비 식당 1 김용세, 김병선
- who? 마이클 잭슨 툰쟁이, 한나나
- 쓸 만한 인간 박정민
- 기분이 태도가 되지 않게 레몬심리
- 용선생 처음 세계사1: 고대 문명~중세: 고대 문명~중세 김선혜, 정지윤, 노남희, 뭉선생, 윤효식, 이우일, 김선빈, 사회평론 역사연구소
- 돈의 속성 김승호
- 바람의 화원 1 이정명
- Henry and Mudge: The First Book Cynthia Rylant
- who? 단군·주몽 정병훈, 최재훈
- who? 이순신 스튜디오청비, 이수겸
- 죽이고 싶은 아이 1 이꽃님
- 무작정 쇼트트랙 이재영
- 어린 왕자 앙투안 드 생텍쥐페리
언제 어디서나 스토리텔
국내 유일 해리포터 시리즈 오디오북
5만권이상의 영어/한국어 오디오북
키즈 모드(어린이 안전 환경)
월정액 무제한 청취
언제든 취소 및 해지 가능
오프라인 액세스를 위한 도서 다운로드
스토리텔 언리미티드
5만권 이상의 영어, 한국어 오디오북을 무제한 들어보세요
13800 원 /월
계정 1개
무제한 청취
사용자 1인
무제한 청취
언제든 해지하실 수 있어요
패밀리
친구 또는 가족과 함께 오디오북을 즐기고 싶은 분들을 위해
매달 21500 원 원 부터
2-3 개 계정
무제한 청취
2-3 계정
무제한 청취
언제든 해지하실 수 있어요
본인 + 1 가족 구성원
2 개 계정21500 원 /월
