Satoshi Munakata, Takaki Nakamura
This paper addresses the challenge of efficiently operating unstructured data stored in object storage using table formats. Table formats enable data analysts to query objects within object storage buckets like a database, which facilitates efficient data retrieval in data science workloads. Conventional table formats require converting unstructured data into structured ones such as Parquet or ORC formats and storing them in separate buckets. Conventional data conversion requires extra time and effort, and separate bucket storage leads to duplicating data in data management manner. To overcome this problem, we propose a new table format that introduces virtual schema structure to unstructured data without structured transformation by leveraging object metadata. Our method allows a bucket containing unstructured data to be directly treated as a table and makes it possible to eliminate conversion overhead and simplify data management. We evaluate query execution time using TPCx-BB benchmark data. The results show that the proposed table format achieves query performance equal to or better than that of conventional ones.