mamlr

edevries

mamlr

Archived

Author	SHA1	Message	Date
Erik de Vries	8d19333e59	query_gen_actors: changed script order for belgium exceptions	6 years ago
Erik de Vries	3bfe61e425	query_gen_actors: fixed implementation of Belgian exceptions	6 years ago
Erik de Vries	81697345cb	modelizer: removed breaking code	6 years ago
Erik de Vries	9ca952ca89	elastic_update: removed wait_for from url	6 years ago
Erik de Vries	8051a81b66	actorizer, dfm_gen, modelizer, out_parser: replaced all instances of detectCores by cores parameter (which defaults to detectCores)	6 years ago
Erik de Vries	ac37d836f5	elasticizer: added scroll_clear to null hits as well	6 years ago
Erik de Vries	75623856f7	elasticizer: updated scroll_clear to use conn object	6 years ago
Erik de Vries	c2d666c81d	bogus commit	6 years ago
Erik de Vries	e34460bf0f	elasticizer: clear scroll context when finishing query	6 years ago
Erik de Vries	9bd526fee0	elasticizer: fixed compatibility issues with elastic v1.0.0	6 years ago
Erik de Vries	f2312f65d5	elasticizer: update to account for syntax change in newer package versions	6 years ago
Erik de Vries	f6006eb9ba	actorizer: simplified pre/postfix check, only for NA, replace empty strings by NA beforehand	6 years ago
Erik de Vries	298099a4e6	actorizer: fix to deal with empty updates (ie dont do an update)	6 years ago
Erik de Vries	6961c0b866	query_gen_actors: updated actorid filter to use the keyword subfield	6 years ago
Erik de Vries	703b5e59a4	actorizer: fixed exceptionizer by adding whitespace before and after sentence, which is necessary because of negative regex (match anything before or after the highlight string that is NOT x actually requires something to be in front or after)	6 years ago
Erik de Vries	593d2de6e2	actorizer: add pre_tags and post_tags to argument list bulk_writer: updated to use _doc doctype query_gen_actors: added NA for all searches that don't have pre- or postfixes	6 years ago
Erik de Vries	a1b6c6a7cb	actorizer, query_gen_actors: revamped actor searches entirely elasticizer: updated script for use with ES 7.x	6 years ago
Erik de Vries	88fc4ec53c	dfm_gen: changed out_parser call to mamlr:::out_parser	6 years ago
Erik de Vries	90fdbcc982	out_parser: parallelized when not in windoze	6 years ago
Erik de Vries	6414f759bd	actorizer: parallelized calculation of marker positions	6 years ago
Erik de Vries	522c872dba	out_parser: moved cleaning regex to end of pipeline, to prevent collissions with other (mandatory) regex cleaning	6 years ago
Erik de Vries	5b9793cd8c	actorizer: removed nested mclapply	6 years ago
Erik de Vries	1a4ba19546	actorizer: Removed udmodel dependencies, commented code, changed nested lists to flat lists bulk_writer: changed handling of single-row dataframe parsing to JSON elastic_update: changed function to return instead of print appData on error ud_update: Changed nested lists to flat lists, and added start and end character positions	6 years ago
Erik de Vries	3abc3056e0	actorizer: fix to columns selected for actors variable, removed udmodel requirement	6 years ago
Erik de Vries	41c86ea116	actorizer, ud_update: Updated ud parsing and actorizer to work based on character positions. This code is used for local testing	6 years ago
Erik de Vries	eae1a22609	actorizer: update to use '\|\|\|' as highlight indicator, and set up ud output merging accordingly	6 years ago
Erik de Vries	5665b6d622	actorizer: more fixes to punctuation	6 years ago
Erik de Vries	cd05733648	actorizer: Additional fix for missing punctuation (see previous commit)	6 years ago
Erik de Vries	09732a1b5a	actorizer: quick fix for problem where original UK UD output does not have a dot at the end of the document, but the actor output does (old vs new parsing)	6 years ago
Erik de Vries	835d2332bc	actorizer: now uses the original udpipe output for sentence and token ids. When the actorized and original udpipe output do not have the same number of rows, it prints an error and sets err to TRUE in actorDetails	6 years ago
Erik de Vries	e70b6ccf7a	actorizer: fixed sentence_count and out_parser calls out_parser: Added comment with old regex	6 years ago
Erik de Vries	9b0ac775af	class_update: add ver variable to set version for class updated articles	7 years ago
Erik de Vries	85306007f4	class_update: added words and clean parameters, in addition to text parameter, to be able to set data preprocessing exactly the same as in the trained model	7 years ago
Erik de Vries	e110780ad5	merger: idiotic fix for a non-problem, see comment on line 32	7 years ago
Erik de Vries	ce5f812252	dfm_gen, merger: Added option for generating lemma_upos hybrids for merged field merger: Added custom clean option (sometimes not cleaning is preferred, even with lemmas) merger, out_parser: Updated regex for filtering out non-words to also include email addresses (containing both @ and .)	7 years ago
Erik de Vries	386ac42aee	lemma_writer: new function to write raw lemma's (without interpunction) to text file. Is structured as elasticizer update function (despite not updating anything on the server)	7 years ago
Erik de Vries	4407a99774	actorizer: fix to get actual number of sentence occurences of actor	7 years ago
Erik de Vries	96e869fa6b	actorizer: previous commit was wrong, only add is an option, removed type variable	7 years ago
Erik de Vries	98219c807c	actorizer: Added type option, to choose between setting or adding to the actor variables, defaults to add (should normally not be changed)	7 years ago
Erik de Vries	e3b57ed9e3	actorizer: added clean = F to have the exact same behavior in ud_update and actorizer	7 years ago
Erik de Vries	7218f6b8d0	dupe_detect: fixed error on no duplicates	7 years ago
Erik de Vries	b9be372543	dupe_detect: fix to get correct colnames from simil (disable stringsAsFactors and convert col values to numeric)	7 years ago
Erik de Vries	1955692346	dfm_gen, out_parser: updated documentation dupe_detect: major fix to function, no longer using rownames for article ids	7 years ago
Erik de Vries	34531b0da8	out_parser: added option to clean output using regex to remove numbers and non-words dfm_gen, ud_update: updated functions to make use of out_parser cleaning option merger: updated regex for cleaning lemmatized output	7 years ago
Erik de Vries	5851c56369	query_string: updated check for fields value	7 years ago
Erik de Vries	4f8b1f2024	elasticizer: renamed size parameter to batch_size, created max_batch parameter to limit the number of results returned query_string: renamed x parameter to query, added fields parameter to select what fields to return and random boolean parameter to define whether the returned results should be randomized	7 years ago
Erik de Vries	d0e9bf565b	dupe_detect: Reset the _delete value to 1 out_parser: fix to sentence parsing, add additional (empty) string at end of merged field, to make merged field end on .	7 years ago
Erik de Vries	ea8cfb071f	dupe_detect: updated _delete var to be 2 when delete is true	7 years ago
Erik de Vries	0a3bdb630b	actorizer, dfm_gen, ud_update: unified output parsing from _source and highlight fields into a single function (out_parser) out_parser: function to parse raw text output into a single field, either from _source or highlight fields dupe_detect: updated function to use 'ver' parameter for versioning	7 years ago
Erik de Vries	9e5a1e3354	ud_update: removed mc.preschedule = F	7 years ago

1 2

97 Commits (8d19333e59a6c2968b16b8b83c9206330058d7c7)