edevries/mamlr - mamlr - git.thijsdevries.net

edevries

/

mamlr

Archived

Author	SHA1	Message	Date
Your Name	4b4d860235	class_update: remove dfm_gen multicore option dfm_gen: remove multicore, update merger() code elasticizer: changed filenaming scheme for dump option merger: Fixed bug where an NA lemma would cause the entire document to become NA. Now the NA lemmas are filtered out before merging ud_update: removed parallel processing, changed script to save bulk updates in .Rds files instead of sending them straight away	5 years ago
Your Name	9eae486a80	separated data preprocessing routines class_update: check if there are idf values associated with model, before applying weights estimator: make use of preproc() function for data preprocessing preproc: function containing all logic with regards to text data preprocessing and weighting	5 years ago
Your Name	a01a53f105	class_update: added cores parameter for multicore processing of sources when using lemmas	5 years ago
Your Name	d9f936c566	modelizer: tf-idf application updated, final model now also includes idf values from training set, explicitly setting positive category in binary classification for confusion matrices, minor code fixes dfm_gen: added old junk codes for recoding, and removed deprecated ngrams parameter from dfm function class_update: removed dfm_words parameter, which is replaced by the force = T parameter in predict(), training/model idf is now applied to unseen data DESCRIPTION: added quanteda.textmodels as new dependency, since these have been separated from base quanteda 2.0.0 onwards	5 years ago
Erik de Vries	9b0ac775af	class_update: add ver variable to set version for class updated articles	7 years ago
Erik de Vries	85306007f4	class_update: added words and clean parameters, in addition to text parameter, to be able to set data preprocessing exactly the same as in the trained model	7 years ago
Erik de Vries	9f3418ef37	class_update; dfm_gen; merger: updated functions to accept text parameter for both old style 'lemmas' and new style 'ud'	7 years ago
Erik de Vries	0e8c127b86	bulk_writer: fixes for JSON generation and added exception for use of 'tokens' varname class_update/elastic_update: Moved response checking to elastic_update dupe_detect: Finalized dupe_detect	7 years ago
Erik de Vries	6bb8f9b635	class_update: added explicit httr::: references	7 years ago
Erik de Vries	f543d658bd	Major overhaul to ES bulk update integration. Added support for both setting and appending to variables	7 years ago
Erik de Vries	d203de0b2a	Updated elasticizer docs, created modelizer and class_update functions	7 years ago