R: A Completion Generator for R
Plain - MapLibre Tile Specification
MapLibre Tile Specification
Checks Session Status
Provides tools to check variables contained in the user environment, and inspect the currently loaded package namespaces. The intended use is to allow user scripts to throw errors or warnings if unwanted variables exist or if unwanted packages are loaded.
Advanced and Fast Data Transformation in R
A large C/C++-based package for advanced data transformation and statistical computing in R that is extremely fast, class-agnostic, robust, and programmer friendly. Core functionality includes a rich set of S3 generic grouped and weighted statistical functions for vectors, matrices and data frames, which provide efficient low-level vectorizations, OpenMP multithreading, and skip missing values by default. These are integrated with fast grouping and ordering algorithms (also callable from C), and efficient data manipulation functions. The package also provides a flexible and rigorous approach to time series and panel data in R, fast functions for data transformation and common statistical procedures, detailed (grouped, weighted) summary statistics, powerful tools to work with nested data, fast data object conversions, functions for memory efficient R programming, and helpers to effectively deal with variable labels, attributes, and missing data. It seamlessly supports base R objects/classes as well as units, integer64, xts/ zoo, tibble, grouped_df, data.table, sf, and pseries/pdata.frame.
Comparing R's {targets} and dbt for Data Engineering
I’m getting more and more into data engineering these days and having used R for a long time, I’m seeing a lot of problems that look nail-shaped to my R-shaped hammer. The available tools to solve those problems exist for (presumably) very good reasons, so I wanted to take some time to dig into how to use them and compare their workflows to what I would otherwise naively do in R.
Dagster Pipes Protocol for R
Implements the Dagster Pipes protocol, enabling R scripts to communicate with the Dagster orchestrator. R scripts can receive execution context and report asset materializations, check results, and log messages back to Dagster.
Broadcasting: Scalars or vectors | Josiah Parry
Apache Arrow, Rust, and cross-langauge data science | Josiah Parry
geodesy/ruminations/000-rumination.md at main · busstoptaktik/geodesy
Rust geodesy
sx
sx: Scalable Spatial Data Analysis
RFC list — GDAL documentation
OutputStream classes — OutputStream
FileOutputStream is for writing to a file;
BufferOutputStream writes to a buffer;
You can create one and pass it to any of the table writers, for example.
use geoarrow · Issue #2 · ropensci/geotargets
See examples from @anthonynorth njtierney/demo-geotargets#1
README
Read geometry vectors — wk_handle.wk_crc
The handler is the basic building block of the wk package. In
particular, the wk_handle() generic allows operations written
as handlers to "just work" with many different input types. The
wk package provides the wk_void() handler, the wk_format()
handler, the wk_debug() handler, the wk_problems() handler,
and wk_writer()s for wkb(), wkt(), xy(), and sf::st_sfc())
vectors.
paleolimbot/gpkg: Proof of Concept 'GeoPackage' to Arrow Converter
Proof of Concept 'GeoPackage' to Arrow Converter
AURIN-OFFICE/geoparquet_data_manager: When facing large geopackages, you can chunk it in geoparquet files and process each chunk seperately.
When facing large geopackages, you can chunk it in geoparquet files and process each chunk seperately. - AURIN-OFFICE/geoparquet_data_manager
Sqlite Database Pragma Usage - Cross-platform Rust Components
PRAGMA
s
06 - Being PRAGMAtic with SQLite
I have been too used to PostgreSQL in the time I have spent in my career so far. PostgreSQL and MySQL are what I had worked with in the beginning.
https://gdal.org/en/stable/programs/gdal_vector_export_schema.html
gdal driver parquet create-metadata-file — GDAL documentation
Function Argument Validation
Validate function arguments succinctly with informative error messages and optional automatic type casting and size recycling. Enable schema-based assertions by attaching reusable rules to data.frame and list objects for use throughout workflows.
flowr — Easy, scalable big data pipelines using HPCC (high performance computing cluster)
Easy, scalable big data pipelines using HPCC (high performance computing cluster)
pygeoapi-prefect
pygeoapi process manager powered by prefect
Orchestrating AI-driven Geospatial Workflows with Prefect
Explore how orchestrating AI workflows with Prefect can streamline complex geospatial tasks. Learn why Prefect is key for orchestrating AI workflows.
Pragma statements supported by SQLite
PRAGMA schema.cache_size;
PRAGMA schema.cache_size = pages;
PRAGMA schema.cache_size = -kibibytes;
Query or change the suggested maximum number of database disk pages that SQLite will hold in memory at once per open database file. Whether or not this suggestion is honored is at the discretion of the Application Defined Page Cache. The default page cache that is built into SQLite honors the request, however alternative application-defined page cache implementations may choose to interpret the suggested cache size in different ways or to ignore it altogether. The default suggested cache size is -2000, which means the cache size is limited to 2048000 bytes of memory. The default suggested cache size can be altered using the SQLITE_DEFAULT_CACHE_SIZE compile-time options. The TEMP database has a default suggested cache size of 0 pages.
If the argument N is positive then the suggested cache size is set to N. If the argument N is negative, then the number of cache pages is adjusted to be a number of pages that would use approximately abs(N*1024) bytes of memory based on the current page size. SQLite remembers the number of pages in the page cache, not the amount of memory used. So if you set the cache size using a negative number and subsequently change the page size (using the PRAGMA page_size command) then the maximum amount of cache memory will go up or down in proportion to the change in page size.
Backwards compatibility note: The behavior of cache_size with a negative N was different prior to version 3.7.10 (2012-01-16). In earlier versions, the number of pages in the cache was set to the absolute value of N.
When you change the cache size using the cache_size pragma, the change only endures for the current session. The cache size reverts to the default value when the database is closed and reopened.
The default page cache implemention does not allocate the full amount of cache memory all at once. Cache memory is allocated in smaller chunks on an as-needed basis. The page_cache setting is a (suggested) upper bound on the amount of memory that the cache can use, not the amount of memory it will use all of the time. This is the behavior of the default page cache implementation, but an application defined page cache is free to behave differently if it wants.
Targeting database tables in workflows
My work on the Department of Ecology’s Safety of Oil Transportation Act risk model has been an opportunity for me to explore some of the newer tools available in R for reproducible workflows. Taking the time to learn and implement these tools has been incredibly helpful, both because the model requirements were still being nailed down while I was developing it (and thus I needed to be able to easily re-run things and identify changes to results) and because the sheer volume of data requires we use parallel processing approaches in order to achieve feasible run times. I identified the targets package as an excellent tool to achieve both of these requirements, as it not only provides a framework for running and tracking analysis pipelines (which I use for ETL procedures and scheduling model runs) but also allows us to seamlessly switch to parallel approaches using future and backends such as future.callr or future.batchtools.
gdal driver gpkg validate — GDAL documentation
OGR SQL dialect and SQLITE SQL dialect — GDAL documentation
OGR SQL dialect and SQLITE SQL dialect
The GDALDataset supports executing commands against a datasource via the GDALDataset::ExecuteSQL() method. How such commands are evaluated is dependent on the datasets.
For most file formats (e.g. Shapefiles, GeoJSON, MapInfo files), the built-in OGR SQL dialect dialect will be used by defaults. It is also possible to request the SQL SQLite dialect alternate dialect to be used, which will use the SQLite engine to evaluate commands on GDAL datasets.
All OGR drivers for database systems: MySQL, PostgreSQL / PostGIS, Oracle Spatial, SQLite / Spatialite RDBMS, GPKG -- GeoPackage vector, ODBC RDBMS, ESRI Personal GeoDatabase, SAP HANA and MSSQLSpatial - Microsoft SQL Server Spatial Database, override the GDALDataset::ExecuteSQL() function with dedicated implementation and, by default, pass the SQL statements directly to the underlying RDBMS. In these cases the SQL syntax varies in some particulars from OGR SQL. Also, anything possible in SQL can then be accomplished for these particular databases. Generally, only the result of SELECT statements will be returned as layers. For those drivers, it is also possible to explicitly request the OGRSQL and SQLITE dialects, although performance will generally be much less as the native SQL engine of those database systems.
SQL is executed against an GDALDataset, not against a specific layer